Network IaC, Part 3: Deployment and Drift Detection
Deploying reviewed network changes from a separate repository whose runner is the only CI runner that can read the device credentials, with a removal guard, a convergence gate, and a scheduled job that tells drift apart from a change that simply hasn't been deployed yet. Including the bugs a review found in the first version.
on this page
Network IaC from a brownfield lab. Part 3 of a series. Part 1 built the Ansible workflow: collect, model, diff offline, apply, re-collect. Part 2 put it under CI for a team, without ever giving CI a device password. This part deploys what the team approved, and keeps checking afterwards that the devices still match.
Part 2 ended with an approved, merged change that no device had: a second NTP server, 192.168.100.11, on all three devices. CI had shown the exact commands for each vendor. Nothing had sent them.
Sending them is the easy part. apply.yml from part 1 already does it. The hard part is deciding where the device credentials can live so that only reviewed code can use them, in a system where any engineer can push a branch with their own workflow file. The second hard part is what happens after the deployment: someone changes a device by hand, a device reboots, or a reviewed change is merged and waits for its deployment. A scheduled job has to tell those apart, and only one of them is an incident.
I built this, tested it, and wrote it up. Then a review of the write-up found three bugs in the workflows. All three are fixed. The main one is tested on the same lab, the other two only in an offline simulation. This post covers both versions, because the main bug was in the very thing the drift job exists for.
Lab setup#
- Same lab and the same model as parts 1 and 2: Arista vEOS, Cisco Nexus 9300v and Cumulus VX in EVE-NG. The model owns NTP and syslog server lines, nothing else.
- Gitea 1.27.1 in Docker, organization
netops, Gitea Actions withact_runner0.2.11 in host mode. - One runner host, two runners. The CI runner from part 2 (Unix user
act-runner, labellab, registered to the organization) and a new deploy runner (Unix useract-deploy, labeldeploy, registered to one repository). - The “team” is still partly simulated.
darrenis the only owner,reviewer01is the second engineer. I control both. - This part writes to devices. Every deployment below happened, and every log line quoted is from a real run, trimmed to the relevant lines. Times are UTC, all from the same day.
Where the device credentials can live#
Part 2 established the rule: unreviewed code (any branch) may read the backups, but must never get device credentials. Reviewed code (main) may deploy.
Gitea secrets can’t enforce that. A secret is available to every workflow that runs in its repository or organization, whatever branch it was pushed from. Anyone who can push a branch can add a workflow file that prints it. Gitea 1.27 has no way to tie a secret to a branch.
So the vault key isn’t a Gitea secret at all. It’s a file on the runner host, and the question becomes which jobs can run as the one Unix user that can read it:
| Piece | What it is | Who can change it |
|---|---|---|
netops/iac-deploy | a separate, private repository holding only two workflows: deploy and drift | owners only |
| deploy runner | registered to iac-deploy alone, label deploy, runs as act-deploy | the runner host’s admin |
| vault key | /etc/iac/vault_pass, owned by act-deploy, in a 0750 root:act-deploy directory | the key holder |
| Git access | two SSH deploy keys in act-deploy’s home directory: read-only on iac-model, write on iac-backups | owners |
| model code | always cloned at main, never a branch | through reviewed pull requests |
act-runner (CI) and act-deploy (deployment) are separate system users with 0750 home directories, and neither is in the other’s group. A CI job can’t read the key, the deploy keys, or the deploy runner’s own registration.
The runner is the part that does the work. A runner registered to the organization serves every repository in it. A runner registered to a repository serves only that repository. So a branch in iac-model can ask for runs-on: deploy as often as it likes, and Gitea has nowhere to send the job.
Who can deploy#
Gitea teams decide who can start a deployment:
| Team | iac-deploy | Why |
|---|---|---|
deployers (darren) | Code Read, Actions Write | Actions Write is for starting a run by hand. Code Read means a deployer can’t change what the workflow does |
readers | Read | everyone can see every deployment and its log |
engineers | none | their team has Write on iac-model, and a Gitea team has one permission level for all its repositories. Adding iac-deploy would let engineers edit the deploy workflows |
iac-deploy’s main accepts pushes only from owners, force pushes are disabled, and administrators are bound by the rule. In the lab that’s one person. A real team would want two owners, so that no single person controls the path from approval to devices.
The backups get a rule of their own. iac-backups main allows pushes from nobody, with one exception: deploy keys with write access. So from part 3 on, the history of the backups is meant to be written only by the pipeline. Every commit quoted below is authored by iac-pipeline.
The deploy workflow#
It’s started by hand, from the Actions page, with two inputs: which device (or all), and whether the plan may remove lines from devices. It runs five steps:
- Clone the model at
mainwith the read-only deploy key, and the backups with the write key. - Pre-deploy collection. Log in, collect the running config, and commit it if anything changed since the last collection. The plan must be computed against what the devices hold now, not against an old snapshot.
- Plan and guardrails. Run the offline diff from part 1, print exactly what each device would receive, and decide: nothing to do, go, or refuse.
- Apply, then the convergence gate: collect again, diff again, and fail the run unless every device’s owned lines match the model.
- Record the result in the backups, including after a failed apply.
concurrency: lab-devices puts deploy and drift in one queue, and the runner takes one job at a time. A drift check can’t collect while a deployment is changing a device.
“Match the model” needs a precise meaning, here and in everything below. The model owns NTP and syslog server lines. The diff compares only those: lines the model wants that a device lacks (missing), and server lines a device has that the model doesn’t list (extra). Everything else on the device is outside the check. And a matching config line proves the line is there, not that the device syncs its clock or that the syslog server receives anything.
The guardrail: a removal needs a decision#
Sending the missing lines is the change the reviewer approved. Removing extra lines is different: an extra line might be someone’s emergency fix from last night. So the plan refuses to remove anything unless the person starting the deployment ticks allow_removals.
A hand-added NTP server on Arista01, and a deployment started without the box ticked:
model main: 85c2c99 limit: Arista01 allow_removals: false
devices unchanged since the last collection
== Arista01
no ntp server 192.168.100.99
plan: missing=0 extra=1
REFUSED: the plan removes 1 line(s). A removal can mean someone added something by hand.
Check the lines above. If removing them is intended, run again with allow_removals.The run stopped before the apply step, so nothing was sent. The same deployment with the box ticked removed the line three minutes later, and the gate passed.
A confirmation the pipeline can’t skip by accident#
Part 1’s apply.yml asked for confirmation with Ansible’s pause module before sending anything. Before wiring it into a pipeline, I tested what pause does without a terminal. It doesn’t fail. It prints Not waiting for response to prompt as stdin is not interactive and carries on. The confirmation that protected a person at a keyboard would have been silently skipped by every pipeline run.
Now a person has to type yes, and a pipeline has to say explicitly that it was approved. Anything else is refused before a single line is sent:
- name: Confirm (a person types yes; a pipeline passes pipeline_confirmed=true)
ansible.builtin.pause:
prompt: "Type yes to send the lines above to {{ inventory_hostname }}"
register: confirm
when: not (pipeline_confirmed | default(false) | bool)
- name: Refuse unless confirmed (non-interactive pause continues silently)
ansible.builtin.assert:
that: (pipeline_confirmed | default(false) | bool) or (confirm.user_input | default('') == 'yes')
fail_msg: "Not confirmed - nothing sent."The deploy workflow passes -e pipeline_confirmed=true. The approval it stands for is the pull request in iac-model plus the person who pressed Run workflow.
The first full deployment#
model main: 85c2c99 limit: all allow_removals: false
devices unchanged since the last collection
== Arista01: nothing to send
== Cumulus01
nv set service ntp mgmt server 192.168.100.11
nv config apply -y
nv config save
== NXOS
ntp server 192.168.100.11 use-vrf management
plan: missing=2 extra=0Arista01 had received the change earlier, as a one-device first deployment, so only two devices had anything to do. The apply step sends to one device at a time (serial: 1) and prints the exact lines before each push. Then it collects all three again, and the gate checks the result instead of trusting the push:
ok: [Cumulus01] =>
msg: owned=3 missing=0 extra=0
ok: [Arista01] =>
msg: owned=3 missing=0 extra=0
ok: [NXOS] =>
msg: owned=3 missing=0 extra=0Then the play Convergence gate - fail if anything is missing or extra asserted zero drift on all three, and the record step committed c91bf36 post-deploy 85c2c99 (all, gate: success).
NX-OS printed one warning: To ensure idempotency and correct diff the input configuration lines should be similar to how they appear if present in the running configuration on device. That’s the generic warning from nxos_config. The gate is the reason I don’t have to take the module’s word for it either way: the line either appears in the re-collected config in the form the model expects, or the run fails.
Idempotence#
A second full deployment when nothing differs:
model main: 0fc0e79 limit: all allow_removals: false
devices unchanged since the last collection
== Arista01: nothing to send
== Cumulus01: nothing to send
== NXOS: nothing to send
plan: missing=0 extra=0
already converged - nothing to send
nothing to recordThis repeat deployment sent nothing. With CONVERGED already current and the backups unchanged, it created no commit.
The drift job#
Every 30 minutes, and on demand, the drift job collects all devices, commits whatever changed into the backups, and compares the devices’ owned lines with the model at main. Gitea runs schedules in UTC since 1.23. For a 30-minute interval that doesn’t matter.
It’s tempting to fail whenever the devices don’t match main. That’s wrong in a team. If a pull request is merged at 09:00 and deployed at 09:10, the devices don’t match main for ten minutes, and nobody did anything wrong. So the job needs a second reference: a model the devices are known to have matched. If they differ from main and main has moved on since that reference, an approved change is probably waiting for its deployment. If they differ from the very model they matched, something changed them outside the pipeline.
Which reference to use is where the first version went wrong.
The first version#
The first version used a file the deploy workflow writes into the backups after a full deployment:
model 6b90a7a29e12e555275a6e22268104a756473e19
deployed 2026-10-10T09:10:38Z
run http://10.0.20.22:3000/netops/iac-deploy/actions/runs/27That’s DEPLOYED: the model commit the devices were last fully deployed from, and the run that did it. The drift job asked three questions, in this order:
- Do the devices match
main? Then 🟢 green. - If not, is
maindifferent fromDEPLOYED? Then 🟡 pending: the job passes, it’s not an incident. - Otherwise 🔴 drift: the job fails.
All three states showed up on real runs. A second reviewed change, syslog to 192.168.100.11, merged as 6b90a7a and not yet deployed:
== Arista01
logging host 192.168.100.11
== Cumulus01
nv set service syslog mgmt server 192.168.100.11
nv config apply -y
nv config save
== NXOS
logging server 192.168.100.11 use-vrf management
PENDING DEPLOYMENT: model main 6b90a7a is not the deployed model (85c2c99288128e777d3ca202b0537dee480933ec).
Difference: missing=3 extra=0. Not an incident - deploy when ready.And with DEPLOYED at 85c2c99, a hand-added NTP server on Arista01:
committed: changed on Arista01
== Arista01
no ntp server 192.168.100.99
DRIFT: devices differ from the deployed model 85c2c99: missing=0 extra=1.
Someone changed a device outside the pipeline. The backups commit above shows what changed.What the review found#
DEPLOYED is only written when a full deployment sends something and passes the gate. A model change that doesn’t change any device config never produces such a deployment: a README edit, a CI fix, or the backup filter later in this post. After one of those, main and DEPLOYED differ permanently, the devices still match main, and every drift run is green. Then someone changes a device by hand. The devices no longer match main, main differs from DEPLOYED, so the first version says 🟡 pending, passes, and alerts nobody. It keeps saying so on every run, until someone happens to remove the line or the next full deployment moves DEPLOYED.
That’s the drift job failing at the one thing it exists for. This lab had that window too: between merging the backup filter at about 08:47 and the syslog deployment at 09:10, a hand change would have shown as pending.
The review found two more bugs, smaller but real:
- No summary read as “nothing to do”. The plan and the drift job added up
missingandextrafrom whatever per-device summary files existed. No files gave0 0, which reads as converged. In practice an unmatchedlimitis already stopped by Ansible (no hosts to target), but nothing checked that every device had produced its summary. - The record after a failed apply wasn’t fresh. The record step committed whatever backup files existed. If the apply stopped before its own re-collection, those files could hold the state from before the change, under a commit called
post-deploy.
The fix: verified, not deployed#
The fix adds a second file, CONVERGED: the last model all devices were verified against. Three things write it: a full deployment that passes the gate, a full deployment that finds nothing to do, and a green drift run. The drift job now compares main with CONVERGED (falling back to DEPLOYED until the first CONVERGED exists). DEPLOYED stays what it was: the record of the last full deployment that changed devices.
Both other bugs got direct fixes. The plan and the drift job now ask the inventory which devices are selected and require one summary per device, or the step fails. And if the apply didn’t succeed, the record step collects again before committing.
The fix went in as a pull request in iac-deploy, merged by the owner. With one owner, nobody else reviewed it, which is the limitation from earlier showing up in practice.
The main fix, tested on the same lab#
The test was the exact failure case: a model change that doesn’t touch any device, then a hand change.
The model change was a README line, merged in iac-model as 0fc0e79 through the usual reviewed pull request. Then ntp server 192.168.100.99 on Arista01 again. The first drift run on the new code still said pending:
== Arista01
no ntp server 192.168.100.99
PENDING DEPLOYMENT: model main 0fc0e79 differs from the last verified model (6b90a7a29e12e555275a6e22268104a756473e19).
Difference: missing=0 extra=1. Not an incident - deploy when ready.That’s correct at that moment. No CONVERGED existed yet, and 0fc0e79 had never been verified against the devices, so the job couldn’t know whether the extra line was a hand change or part of an undeployed change. After the hand change was removed, the next drift run was green and recorded the verification in the backups:
d7e35ea verified 0fc0e79 (drift-check, all devices)
model 0fc0e79eec47f39cd07c2e88758b123c7eaea407
verified 2026-10-10T14:44:01Z
run http://10.0.20.22:3000/netops/iac-deploy/actions/runs/44Then the hand change again, and this time I left it to the schedule. The 15:00 scheduled run:
committed: changed on Arista01
== Arista01
no ntp server 192.168.100.99
DRIFT: devices differ from model 0fc0e79, which they were last verified against: missing=0 extra=1.
Someone changed a device outside the pipeline. The backups commit above shows what changed.Red, from a run nobody started by hand, with DEPLOYED still at 6b90a7a. The first version would have called the same situation pending.
The other two fixes I could only test offline. A simulation with fake devices fails the step when one device’s summary is missing, and collects again when the apply fails. Breaking a real deployment on purpose wasn’t worth it in this lab.
One limitation remains even with the fix. After a model change, drift detection is only armed again once a green drift run or a full deployment has verified the devices against the new main. A hand change made before that shows as pending and alerts nobody, on every run. The window stays open until the devices match the new model and a run records that verification.
Red has two causes#
A drift run also fails when it can’t look. Twice in this part, a device was unreachable:
fatal: [NXOS]: FAILED! =>
changed: false
msg: timed outfatal: [Cumulus01]: UNREACHABLE! =>
changed: false
msg: 'Task failed: Failed to connect to the host via ssh: ssh: connect to host 192.168.100.13
port 22: Connection timed out'
unreachable: trueIn both cases the collection step failed, nothing was committed, and the classification never ran. That’s deliberate: “I couldn’t check” must never come out as “no drift”. But the run is red either way, so a red run means “look at the log”, which says whether it was a collection failure or configuration drift.
A device reboot is not a config change#
The first drift commit of the day wasn’t drift. After NXOS came back from the outage above, the drift job committed e79e918 drift-check: changed on NXOS, and the whole diff was one line:
+!No configuration change since last restart
NX-OS adds that comment to the top of the running config after a reload. Part 1’s collection already stripped two other NX-OS header lines that change on their own, !Time: and !Running configuration last done at:, so that the backups’ history only records real changes. This third one wasn’t in the filter. It never affected the classification, because it isn’t an owned line. But it put a commit into the audit trail that said “changed” when nothing had.
The fix went through the normal process, a pull request in iac-model with CI and a review:
volatile_headers: '(?m)^!((Time|Running configuration last done at):|No configuration change since last restart).*\n'The colon stays required for the first two, so an unrelated comment that happens to start with !Time isn’t removed. NX-OS dropped the line itself at the next config change: the full deployment’s backups commit, c91bf36, removes it along with adding the NTP server. So the filter hasn’t had to remove it yet. It will after the next reload.
Notification#
Gitea can mail the user who started a failed run. In this lab it first didn’t, for two reasons:
- Mail was disabled (
Mailer Enabled: false). In the lab, Gitea now sends to a mail catcher (Mailpit) running next to it, which keeps every mail and delivers none, so nothing leaves the lab network. - Actions mails need their own switch. Mail for workflow runs arrived in Gitea 1.25 (go-gitea/gitea#34982), and in this lab it only sent with
ENABLE_NOTIFY_MAIL = truein the[service]section. A working test mail from the admin page didn’t prove that part.
With both in place, the two failed drift runs I had started by hand each produced one mail, Run failed: drift.yml (11352e5317). One was the Cumulus01 outage, one was drift. The subject is the same for both, and the commit in it is the iac-deploy workflow’s, not the model’s. Green and pending runs, and the deployments, sent nothing.
The scheduled run that caught the hand change at 15:00 sent no mail at all. The discussion on that Gitea change notes that a scheduled run has no real user to send it to. So in this setup, the job nobody watches records drift and turns red, but nobody is told. A real setup needs its own alert, for example a final workflow step that calls the team’s chat webhook when the job fails. The lab has no chat to send to, so this one stays open.
Testing the boundary#
Two more tests check the boundary itself.
A branch can’t use the deploy runner. I pushed a branch to iac-model with this workflow, which is exactly the attack the design is meant to stop:
name: steal
on: push
jobs:
grab:
runs-on: deploy
steps:
- run: cat /etc/iac/vault_passThe same push started the normal CI workflow, which the organization runner picked up and finished in 1 minute 46 seconds. The steal run stayed at Waiting, with a total duration of 0 seconds, for over an hour, until I cancelled it. No runner reachable from iac-model has the deploy label, so Gitea never handed the job out. Nothing about this depends on someone noticing the workflow in review.
The second engineer can’t deploy. Signed in as reviewer01, the deploy and drift workflows show their full run history, but neither shows a Run workflow button. The team can see every deployment and its log. Starting one is reserved for deployers.
The workflows#
Both files, exactly as they are on iac-deploy main after the fix:
name: deploy
on:
workflow_dispatch:
inputs:
limit:
description: "Device to deploy to, or all"
required: true
default: Arista01
type: choice
options: [Arista01, NXOS, Cumulus01, all]
allow_removals:
description: "Allow the plan to remove lines from devices"
required: true
default: false
type: boolean
# One collection or deployment against the lab at a time, across deploy and drift.
concurrency:
group: lab-devices
cancel-in-progress: false
env:
# Device credentials: a file only act-deploy can read. Never a Gitea secret.
ANSIBLE_VAULT_PASSWORD_FILE: /etc/iac/vault_pass
IAC_BACKUP_DIR: ${{ github.workspace }}/backups
MODEL_SSH: ssh -i /var/lib/act-deploy/.ssh/model_ro -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222
BACKUPS_SSH: ssh -i /var/lib/act-deploy/.ssh/backups_rw -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222
LIMIT: ${{ github.event.inputs.limit }}
ALLOW_REMOVALS: ${{ github.event.inputs.allow_removals }}
jobs:
deploy:
runs-on: deploy
steps:
- name: Model at main (reviewed code only), and the backups
run: |
rm -rf model backups
GIT_SSH_COMMAND="$MODEL_SSH" git clone -q --depth 1 --branch main [email protected]:netops/iac-model.git model
GIT_SSH_COMMAND="$BACKUPS_SSH" git clone -q --branch main [email protected]:netops/iac-backups.git backups
git -C backups config user.name "iac-pipeline"
git -C backups config user.email "[email protected]"
echo "model main: $(git -C model rev-parse --short HEAD) limit: $LIMIT allow_removals: $ALLOW_REMOVALS"
- name: Toolchain (locked)
working-directory: model
run: uv sync --locked && mkdir -p logs
- name: Pre-deploy collection
working-directory: model
run: |
uv run --locked ansible-playbook playbooks/backup.yml --limit "$LIMIT"
sha=$(git rev-parse --short HEAD)
git -C "$IAC_BACKUP_DIR" add -A
git -C "$IAC_BACKUP_DIR" diff --cached --quiet && echo "devices unchanged since the last collection" || {
git -C "$IAC_BACKUP_DIR" commit -q -m "pre-deploy $sha ($LIMIT)"
GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
echo "devices had changed since the last collection - committed as pre-deploy $sha"; }
- name: Plan and guardrails
id: plan
working-directory: model
run: |
uv run --locked ansible-playbook playbooks/diff.yml --limit "$LIMIT"
for f in remediation/*.cfg; do
if [ -s "$f" ]; then echo "== $(basename "$f" .cfg)"; cat "$f"; else echo "== $(basename "$f" .cfg): nothing to send"; fi
done
# One summary per selected device, or the step fails. No summary must never read as "nothing to do".
hosts=$(uv run --locked ansible eos,nxos,cumulus --list-hosts --limit "$LIMIT")
hosts=$(tail -n +2 <<< "$hosts")
test -n "$hosts" || { echo "no devices match '$LIMIT'"; exit 1; }
counts=$(python3 -c 'import json,sys
m=e=0
for h in sys.argv[1:]:
d=json.load(open("remediation/%s.json" % h)); m+=int(d["missing"]); e+=int(d["extra"])
print(m, e)' $hosts)
read missing extra <<< "$counts"
echo "plan: missing=$missing extra=$extra"
if [ "$missing" = 0 ] && [ "$extra" = 0 ]; then echo "already converged - nothing to send"; echo "converged=true" >> "$GITHUB_OUTPUT"; exit 0; fi
if [ "$extra" != 0 ] && [ "$ALLOW_REMOVALS" != "true" ]; then
echo "REFUSED: the plan removes $extra line(s). A removal can mean someone added something by hand."
echo "Check the lines above. If removing them is intended, run again with allow_removals."
exit 1
fi
echo "converged=false" >> "$GITHUB_OUTPUT"
- name: Apply, re-collect and verify (the convergence gate)
id: apply
if: steps.plan.outputs.converged == 'false'
working-directory: model
run: uv run --locked ansible-playbook playbooks/apply.yml --limit "$LIMIT" -e pipeline_confirmed=true
# Runs whenever the plan reached a decision, including after a failed apply.
# DEPLOYED: the last full deployment that sent something and passed the gate.
# CONVERGED: the last model all devices were verified against (gate, no-op deploy, or a green drift check).
- name: Record the result in the backups
if: always() && steps.plan.outputs.converged != ''
working-directory: model
run: |
sha=$(git rev-parse --short HEAD)
stamp() { printf 'model %s\n%s %s\nrun %s/%s/actions/runs/%s\n' "$(git rev-parse HEAD)" "$1" "$(date -u +%FT%TZ)" "${{ github.server_url }}" "${{ github.repository }}" "${{ github.run_number }}"; }
last=$(sed -n 's/^model //p' "$IAC_BACKUP_DIR/CONVERGED" 2>/dev/null || true)
if [ "${{ steps.plan.outputs.converged }}" = true ]; then
msg="verified $sha ($LIMIT, already converged)"
if [ "$LIMIT" = all ] && [ "$last" != "$(git rev-parse HEAD)" ]; then stamp verified > "$IAC_BACKUP_DIR/CONVERGED"; fi
else
if [ "${{ steps.apply.outcome }}" != success ]; then
echo "apply did not succeed - collecting again, so the record shows the state after the attempt"
uv run --locked ansible-playbook playbooks/backup.yml --limit "$LIMIT" || echo "post-attempt collection incomplete: an unreachable device keeps its pre-deploy file"
fi
msg="post-deploy $sha ($LIMIT, gate: ${{ steps.apply.outcome }})"
if [ "$LIMIT" = all ] && [ "${{ steps.apply.outcome }}" = success ]; then
stamp deployed > "$IAC_BACKUP_DIR/DEPLOYED"
stamp verified > "$IAC_BACKUP_DIR/CONVERGED"
fi
fi
git -C "$IAC_BACKUP_DIR" add -A
git -C "$IAC_BACKUP_DIR" diff --cached --quiet && { echo "nothing to record"; exit 0; }
git -C "$IAC_BACKUP_DIR" commit -q -m "$msg"
GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
git --no-pager -C "$IAC_BACKUP_DIR" log --oneline -1name: drift
on:
schedule:
- cron: "*/30 * * * *" # Gitea runs schedules in UTC; every 30 min doesn't care
workflow_dispatch:
concurrency:
group: lab-devices
cancel-in-progress: false
env:
ANSIBLE_VAULT_PASSWORD_FILE: /etc/iac/vault_pass
IAC_BACKUP_DIR: ${{ github.workspace }}/backups
MODEL_SSH: ssh -i /var/lib/act-deploy/.ssh/model_ro -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222
BACKUPS_SSH: ssh -i /var/lib/act-deploy/.ssh/backups_rw -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222
jobs:
drift:
runs-on: deploy
steps:
- name: Model at main, and the backups
run: |
rm -rf model backups
GIT_SSH_COMMAND="$MODEL_SSH" git clone -q --depth 1 --branch main [email protected]:netops/iac-model.git model
GIT_SSH_COMMAND="$BACKUPS_SSH" git clone -q --branch main [email protected]:netops/iac-backups.git backups
git -C backups config user.name "iac-pipeline"
git -C backups config user.email "[email protected]"
- name: Toolchain (locked)
working-directory: model
run: uv sync --locked && mkdir -p logs
- name: Collect, and commit only what changed
working-directory: model
run: |
uv run --locked ansible-playbook playbooks/backup.yml
git -C "$IAC_BACKUP_DIR" add -A
if git -C "$IAC_BACKUP_DIR" diff --cached --quiet; then echo "no device changed since the last collection"; else
changed=$(git -C "$IAC_BACKUP_DIR" diff --cached --name-only | tr '\n' ' ')
git -C "$IAC_BACKUP_DIR" commit -q -m "drift-check: changed on $changed"
GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
echo "committed: changed on $changed"
fi
- name: Compare with the model, and classify
working-directory: model
run: |
uv run --locked ansible-playbook playbooks/diff.yml >/dev/null
# One summary per device, or the step fails. No summary must never read as green.
hosts=$(uv run --locked ansible eos,nxos,cumulus --list-hosts)
hosts=$(tail -n +2 <<< "$hosts")
test -n "$hosts" || { echo "no devices in the inventory"; exit 1; }
counts=$(python3 -c 'import json,sys
m=e=0
for h in sys.argv[1:]:
d=json.load(open("remediation/%s.json" % h)); m+=int(d["missing"]); e+=int(d["extra"])
print(m, e)' $hosts)
read missing extra <<< "$counts"
model=$(git rev-parse HEAD)
# The last model all devices were verified against. Before the first CONVERGED file, fall back to DEPLOYED.
verified=$(sed -n 's/^model //p' "$IAC_BACKUP_DIR/CONVERGED" 2>/dev/null || true)
[ -n "$verified" ] || verified=$(sed -n 's/^model //p' "$IAC_BACKUP_DIR/DEPLOYED" 2>/dev/null || true)
for f in remediation/*.cfg; do if [ -s "$f" ]; then echo "== $(basename "$f" .cfg)"; cat "$f"; fi; done
if [ "$missing" = 0 ] && [ "$extra" = 0 ]; then
echo "GREEN: owned lines match model ${model:0:7} on all devices"
if [ "$model" != "$verified" ]; then
printf 'model %s\nverified %s\nrun %s/%s/actions/runs/%s\n' "$model" "$(date -u +%FT%TZ)" "${{ github.server_url }}" "${{ github.repository }}" "${{ github.run_number }}" > "$IAC_BACKUP_DIR/CONVERGED"
git -C "$IAC_BACKUP_DIR" add CONVERGED
git -C "$IAC_BACKUP_DIR" commit -q -m "verified ${model:0:7} (drift-check, all devices)"
GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
echo "recorded: CONVERGED is now ${model:0:7}"
fi
exit 0
fi
if [ "$model" != "$verified" ]; then
echo "PENDING DEPLOYMENT: model main ${model:0:7} differs from the last verified model (${verified:-none recorded yet})."
echo "Difference: missing=$missing extra=$extra. Not an incident - deploy when ready."
exit 0
fi
echo "DRIFT: devices differ from model ${verified:0:7}, which they were last verified against: missing=$missing extra=$extra."
echo "Someone changed a device outside the pipeline. The backups commit above shows what changed."
exit 1The audit trail#
The backups’ history for the day, all written by the pipeline:
8be96eb 15:28 post-deploy 0fc0e79 (Arista01, gate: success)
781a30e 15:01 drift-check: changed on Arista01
d7e35ea 14:44 verified 0fc0e79 (drift-check, all devices)
d6acc0a 14:42 post-deploy 0fc0e79 (Arista01, gate: success)
8d101b4 14:31 drift-check: changed on Arista01
23fe862 09:10 post-deploy 6b90a7a (all, gate: success)
28855c2 08:52 post-deploy ca8cbf4 (Arista01, gate: success)
deb0b67 08:27 drift-check: changed on Arista01
c91bf36 08:06 post-deploy 85c2c99 (all, gate: success)
661f5d4 06:39 post-deploy 85c2c99 (Arista01, gate: success)
9b9aa76 06:30 drift-check: changed on Arista01
e7c0cce 06:25 post-deploy 85c2c99 (Arista01, gate: success)
e79e918 06:10 drift-check: changed on NXOSRead from the bottom. Morning, first version: a reboot that looked like a change, the first deployment to one device, a hand change and its removal, the full NTP deployment, the hand change again (red this time), its removal, and the syslog change deployed to all three. Afternoon: the hand change during the gap (collected at 14:30 by the first version, minutes before the fix was merged), its removal, the first verification of 0fc0e79, the hand change caught by the schedule, and its removal.
What this does not show#
- One owner. The deploy workflows, the branch rules and the runner registration are all controlled by one person, and the workflow fix was merged without a second reviewer. The design needs at least two owners before it means anything in a team.
- The deploy runner can read the key. That’s its job. What protects the key is that only owner-controlled workflows run on that runner, and they only check out
main. Anyone with root on the runner host can read it too, so that host is part of the trust boundary. - The runner isn’t the only copy of the key. The key holder still has one, for running playbooks by hand. It’s the only copy a CI runner can reach.
- The model is trusted once it’s on
main. A reviewed pull request can change a playbook, and the deploy runner will run it with the key. The review in part 2 is the control. It’s a person, not a machine. - Only owned lines are checked. “Converged” means the NTP and syslog server lines match the model. It says nothing about the rest of the device config, and nothing about whether NTP is synchronized or syslog messages arrive.
- Pending can hide drift. While a merged change waits for deployment, or before the first green run after any model change, a hand change shows as 🟡, because the job can’t tell which difference came from where. It stays that way until a run verifies the devices against the new model. The next deployment does surface it: the plan lists the extra line and refuses to remove it without
allow_removals. - Scheduled failures alert nobody. See the notification section above.
- Two fixes tested offline only: one summary per device, and collecting again after a failed apply.
- The
deployerspermissions are only half tested.reviewer01has no Run workflow button, as intended. But the only deployer isdarren, who is also an owner and would see the button anyway. A deployer who isn’t an owner hasn’t been tested. - Deployment is manual. Merging doesn’t deploy. That’s a choice for a lab with one owner, not a limitation of the tools.
- Trust on first use for the Gitea SSH host key, which the deploy runner learned with
ssh-keyscanon the lab network. And plain HTTP between runner and Gitea, as in part 2.
Conclusion#
Both approved changes are on all three devices, and their owned lines match the model at main. For the lines this part deployed, the backups lead from the device back to the deployment run that sent them, the model commit, and the pull request that approved it. The drift job tells a waiting change from a hand change once it has verified the current model. The first version of it didn’t, and a review of this write-up caught that before the post went out.
Next in this series: so far the pipeline manages two services on three devices. The next parts take it to other platforms and other IaC scenarios.