darrenqu.net

Automation

Network IaC, Part 3: Deployment and Drift Detection

5446 words 26 min read

ansibleci

Deploying reviewed network changes from a separate repository whose runner is the only CI runner that can read the device credentials, with a removal guard, a convergence gate, and a scheduled job that tells drift apart from a change that simply hasn't been deployed yet. Including the bugs a review found in the first version.

on this page

Network IaC from a brownfield lab. Part 3 of a series. Part 1 built the Ansible workflow: collect, model, diff offline, apply, re-collect. Part 2 put it under CI for a team, without ever giving CI a device password. This part deploys what the team approved, and keeps checking afterwards that the devices still match.

Part 2 ended with an approved, merged change that no device had: a second NTP server, 192.168.100.11, on all three devices. CI had shown the exact commands for each vendor. Nothing had sent them.

Sending them is the easy part. apply.yml from part 1 already does it. The hard part is deciding where the device credentials can live so that only reviewed code can use them, in a system where any engineer can push a branch with their own workflow file. The second hard part is what happens after the deployment: someone changes a device by hand, a device reboots, or a reviewed change is merged and waits for its deployment. A scheduled job has to tell those apart, and only one of them is an incident.

I built this, tested it, and wrote it up. Then a review of the write-up found three bugs in the workflows. All three are fixed. The main one is tested on the same lab, the other two only in an offline simulation. This post covers both versions, because the main bug was in the very thing the drift job exists for.

Lab setup#

  • Same lab and the same model as parts 1 and 2: Arista vEOS, Cisco Nexus 9300v and Cumulus VX in EVE-NG. The model owns NTP and syslog server lines, nothing else.
  • Gitea 1.27.1 in Docker, organization netops, Gitea Actions with act_runner 0.2.11 in host mode.
  • One runner host, two runners. The CI runner from part 2 (Unix user act-runner, label lab, registered to the organization) and a new deploy runner (Unix user act-deploy, label deploy, registered to one repository).
  • The “team” is still partly simulated. darren is the only owner, reviewer01 is the second engineer. I control both.
  • This part writes to devices. Every deployment below happened, and every log line quoted is from a real run, trimmed to the relevant lines. Times are UTC, all from the same day.

Where the device credentials can live#

Part 2 established the rule: unreviewed code (any branch) may read the backups, but must never get device credentials. Reviewed code (main) may deploy.

Gitea secrets can’t enforce that. A secret is available to every workflow that runs in its repository or organization, whatever branch it was pushed from. Anyone who can push a branch can add a workflow file that prints it. Gitea 1.27 has no way to tie a secret to a branch.

So the vault key isn’t a Gitea secret at all. It’s a file on the runner host, and the question becomes which jobs can run as the one Unix user that can read it:

PieceWhat it isWho can change it
netops/iac-deploya separate, private repository holding only two workflows: deploy and driftowners only
deploy runnerregistered to iac-deploy alone, label deploy, runs as act-deploythe runner host’s admin
vault key/etc/iac/vault_pass, owned by act-deploy, in a 0750 root:act-deploy directorythe key holder
Git accesstwo SSH deploy keys in act-deploy’s home directory: read-only on iac-model, write on iac-backupsowners
model codealways cloned at main, never a branchthrough reviewed pull requests

act-runner (CI) and act-deploy (deployment) are separate system users with 0750 home directories, and neither is in the other’s group. A CI job can’t read the key, the deploy keys, or the deploy runner’s own registration.

The runner is the part that does the work. A runner registered to the organization serves every repository in it. A runner registered to a repository serves only that repository. So a branch in iac-model can ask for runs-on: deploy as often as it likes, and Gitea has nowhere to send the job.

Who can deploy#

Gitea teams decide who can start a deployment:

Teamiac-deployWhy
deployers (darren)Code Read, Actions WriteActions Write is for starting a run by hand. Code Read means a deployer can’t change what the workflow does
readersReadeveryone can see every deployment and its log
engineersnonetheir team has Write on iac-model, and a Gitea team has one permission level for all its repositories. Adding iac-deploy would let engineers edit the deploy workflows

iac-deploy’s main accepts pushes only from owners, force pushes are disabled, and administrators are bound by the rule. In the lab that’s one person. A real team would want two owners, so that no single person controls the path from approval to devices.

The backups get a rule of their own. iac-backups main allows pushes from nobody, with one exception: deploy keys with write access. So from part 3 on, the history of the backups is meant to be written only by the pipeline. Every commit quoted below is authored by iac-pipeline.

The deploy workflow#

It’s started by hand, from the Actions page, with two inputs: which device (or all), and whether the plan may remove lines from devices. It runs five steps:

  1. Clone the model at main with the read-only deploy key, and the backups with the write key.
  2. Pre-deploy collection. Log in, collect the running config, and commit it if anything changed since the last collection. The plan must be computed against what the devices hold now, not against an old snapshot.
  3. Plan and guardrails. Run the offline diff from part 1, print exactly what each device would receive, and decide: nothing to do, go, or refuse.
  4. Apply, then the convergence gate: collect again, diff again, and fail the run unless every device’s owned lines match the model.
  5. Record the result in the backups, including after a failed apply.

concurrency: lab-devices puts deploy and drift in one queue, and the runner takes one job at a time. A drift check can’t collect while a deployment is changing a device.

“Match the model” needs a precise meaning, here and in everything below. The model owns NTP and syslog server lines. The diff compares only those: lines the model wants that a device lacks (missing), and server lines a device has that the model doesn’t list (extra). Everything else on the device is outside the check. And a matching config line proves the line is there, not that the device syncs its clock or that the syslog server receives anything.

The guardrail: a removal needs a decision#

Sending the missing lines is the change the reviewer approved. Removing extra lines is different: an extra line might be someone’s emergency fix from last night. So the plan refuses to remove anything unless the person starting the deployment ticks allow_removals.

A hand-added NTP server on Arista01, and a deployment started without the box ticked:

model main: 85c2c99  limit: Arista01  allow_removals: false
devices unchanged since the last collection
== Arista01
no ntp server 192.168.100.99
plan: missing=0 extra=1
REFUSED: the plan removes 1 line(s). A removal can mean someone added something by hand.
Check the lines above. If removing them is intended, run again with allow_removals.

The run stopped before the apply step, so nothing was sent. The same deployment with the box ticked removed the line three minutes later, and the gate passed.

A confirmation the pipeline can’t skip by accident#

Part 1’s apply.yml asked for confirmation with Ansible’s pause module before sending anything. Before wiring it into a pipeline, I tested what pause does without a terminal. It doesn’t fail. It prints Not waiting for response to prompt as stdin is not interactive and carries on. The confirmation that protected a person at a keyboard would have been silently skipped by every pipeline run.

Now a person has to type yes, and a pipeline has to say explicitly that it was approved. Anything else is refused before a single line is sent:

- name: Confirm (a person types yes; a pipeline passes pipeline_confirmed=true)
  ansible.builtin.pause:
    prompt: "Type yes to send the lines above to {{ inventory_hostname }}"
  register: confirm
  when: not (pipeline_confirmed | default(false) | bool)

- name: Refuse unless confirmed (non-interactive pause continues silently)
  ansible.builtin.assert:
    that: (pipeline_confirmed | default(false) | bool) or (confirm.user_input | default('') == 'yes')
    fail_msg: "Not confirmed - nothing sent."

The deploy workflow passes -e pipeline_confirmed=true. The approval it stands for is the pull request in iac-model plus the person who pressed Run workflow.

The first full deployment#

model main: 85c2c99  limit: all  allow_removals: false
devices unchanged since the last collection
== Arista01: nothing to send
== Cumulus01
nv set service ntp mgmt server 192.168.100.11
nv config apply -y
nv config save
== NXOS
ntp server 192.168.100.11 use-vrf management
plan: missing=2 extra=0

Arista01 had received the change earlier, as a one-device first deployment, so only two devices had anything to do. The apply step sends to one device at a time (serial: 1) and prints the exact lines before each push. Then it collects all three again, and the gate checks the result instead of trusting the push:

ok: [Cumulus01] =>
    msg: owned=3 missing=0 extra=0
ok: [Arista01] =>
    msg: owned=3 missing=0 extra=0
ok: [NXOS] =>
    msg: owned=3 missing=0 extra=0

Then the play Convergence gate - fail if anything is missing or extra asserted zero drift on all three, and the record step committed c91bf36 post-deploy 85c2c99 (all, gate: success).

NX-OS printed one warning: To ensure idempotency and correct diff the input configuration lines should be similar to how they appear if present in the running configuration on device. That’s the generic warning from nxos_config. The gate is the reason I don’t have to take the module’s word for it either way: the line either appears in the re-collected config in the form the model expects, or the run fails.

Idempotence#

A second full deployment when nothing differs:

model main: 0fc0e79  limit: all  allow_removals: false
devices unchanged since the last collection
== Arista01: nothing to send
== Cumulus01: nothing to send
== NXOS: nothing to send
plan: missing=0 extra=0
already converged - nothing to send
nothing to record

This repeat deployment sent nothing. With CONVERGED already current and the backups unchanged, it created no commit.

The drift job#

Every 30 minutes, and on demand, the drift job collects all devices, commits whatever changed into the backups, and compares the devices’ owned lines with the model at main. Gitea runs schedules in UTC since 1.23. For a 30-minute interval that doesn’t matter.

It’s tempting to fail whenever the devices don’t match main. That’s wrong in a team. If a pull request is merged at 09:00 and deployed at 09:10, the devices don’t match main for ten minutes, and nobody did anything wrong. So the job needs a second reference: a model the devices are known to have matched. If they differ from main and main has moved on since that reference, an approved change is probably waiting for its deployment. If they differ from the very model they matched, something changed them outside the pipeline.

Which reference to use is where the first version went wrong.

The first version#

The first version used a file the deploy workflow writes into the backups after a full deployment:

model 6b90a7a29e12e555275a6e22268104a756473e19
deployed 2026-10-10T09:10:38Z
run http://10.0.20.22:3000/netops/iac-deploy/actions/runs/27

That’s DEPLOYED: the model commit the devices were last fully deployed from, and the run that did it. The drift job asked three questions, in this order:

  1. Do the devices match main? Then 🟢 green.
  2. If not, is main different from DEPLOYED? Then 🟡 pending: the job passes, it’s not an incident.
  3. Otherwise 🔴 drift: the job fails.

All three states showed up on real runs. A second reviewed change, syslog to 192.168.100.11, merged as 6b90a7a and not yet deployed:

== Arista01
logging host 192.168.100.11
== Cumulus01
nv set service syslog mgmt server 192.168.100.11
nv config apply -y
nv config save
== NXOS
logging server 192.168.100.11 use-vrf management
PENDING DEPLOYMENT: model main 6b90a7a is not the deployed model (85c2c99288128e777d3ca202b0537dee480933ec).
Difference: missing=3 extra=0. Not an incident - deploy when ready.

And with DEPLOYED at 85c2c99, a hand-added NTP server on Arista01:

committed: changed on Arista01
== Arista01
no ntp server 192.168.100.99
DRIFT: devices differ from the deployed model 85c2c99: missing=0 extra=1.
Someone changed a device outside the pipeline. The backups commit above shows what changed.

What the review found#

DEPLOYED is only written when a full deployment sends something and passes the gate. A model change that doesn’t change any device config never produces such a deployment: a README edit, a CI fix, or the backup filter later in this post. After one of those, main and DEPLOYED differ permanently, the devices still match main, and every drift run is green. Then someone changes a device by hand. The devices no longer match main, main differs from DEPLOYED, so the first version says 🟡 pending, passes, and alerts nobody. It keeps saying so on every run, until someone happens to remove the line or the next full deployment moves DEPLOYED.

That’s the drift job failing at the one thing it exists for. This lab had that window too: between merging the backup filter at about 08:47 and the syslog deployment at 09:10, a hand change would have shown as pending.

The review found two more bugs, smaller but real:

  • No summary read as “nothing to do”. The plan and the drift job added up missing and extra from whatever per-device summary files existed. No files gave 0 0, which reads as converged. In practice an unmatched limit is already stopped by Ansible (no hosts to target), but nothing checked that every device had produced its summary.
  • The record after a failed apply wasn’t fresh. The record step committed whatever backup files existed. If the apply stopped before its own re-collection, those files could hold the state from before the change, under a commit called post-deploy.

The fix: verified, not deployed#

The fix adds a second file, CONVERGED: the last model all devices were verified against. Three things write it: a full deployment that passes the gate, a full deployment that finds nothing to do, and a green drift run. The drift job now compares main with CONVERGED (falling back to DEPLOYED until the first CONVERGED exists). DEPLOYED stays what it was: the record of the last full deployment that changed devices.

Both other bugs got direct fixes. The plan and the drift job now ask the inventory which devices are selected and require one summary per device, or the step fails. And if the apply didn’t succeed, the record step collects again before committing.

The fix went in as a pull request in iac-deploy, merged by the owner. With one owner, nobody else reviewed it, which is the limitation from earlier showing up in practice.

The main fix, tested on the same lab#

The test was the exact failure case: a model change that doesn’t touch any device, then a hand change.

The model change was a README line, merged in iac-model as 0fc0e79 through the usual reviewed pull request. Then ntp server 192.168.100.99 on Arista01 again. The first drift run on the new code still said pending:

== Arista01
no ntp server 192.168.100.99
PENDING DEPLOYMENT: model main 0fc0e79 differs from the last verified model (6b90a7a29e12e555275a6e22268104a756473e19).
Difference: missing=0 extra=1. Not an incident - deploy when ready.

That’s correct at that moment. No CONVERGED existed yet, and 0fc0e79 had never been verified against the devices, so the job couldn’t know whether the extra line was a hand change or part of an undeployed change. After the hand change was removed, the next drift run was green and recorded the verification in the backups:

d7e35ea verified 0fc0e79 (drift-check, all devices)

model 0fc0e79eec47f39cd07c2e88758b123c7eaea407
verified 2026-10-10T14:44:01Z
run http://10.0.20.22:3000/netops/iac-deploy/actions/runs/44

Then the hand change again, and this time I left it to the schedule. The 15:00 scheduled run:

committed: changed on Arista01
== Arista01
no ntp server 192.168.100.99
DRIFT: devices differ from model 0fc0e79, which they were last verified against: missing=0 extra=1.
Someone changed a device outside the pipeline. The backups commit above shows what changed.

Red, from a run nobody started by hand, with DEPLOYED still at 6b90a7a. The first version would have called the same situation pending.

The other two fixes I could only test offline. A simulation with fake devices fails the step when one device’s summary is missing, and collects again when the apply fails. Breaking a real deployment on purpose wasn’t worth it in this lab.

One limitation remains even with the fix. After a model change, drift detection is only armed again once a green drift run or a full deployment has verified the devices against the new main. A hand change made before that shows as pending and alerts nobody, on every run. The window stays open until the devices match the new model and a run records that verification.

Red has two causes#

A drift run also fails when it can’t look. Twice in this part, a device was unreachable:

fatal: [NXOS]: FAILED! =>
    changed: false
    msg: timed out
fatal: [Cumulus01]: UNREACHABLE! =>
    changed: false
    msg: 'Task failed: Failed to connect to the host via ssh: ssh: connect to host 192.168.100.13
        port 22: Connection timed out'
    unreachable: true

In both cases the collection step failed, nothing was committed, and the classification never ran. That’s deliberate: “I couldn’t check” must never come out as “no drift”. But the run is red either way, so a red run means “look at the log”, which says whether it was a collection failure or configuration drift.

A device reboot is not a config change#

The first drift commit of the day wasn’t drift. After NXOS came back from the outage above, the drift job committed e79e918 drift-check: changed on NXOS, and the whole diff was one line:

+!No configuration change since last restart

NX-OS adds that comment to the top of the running config after a reload. Part 1’s collection already stripped two other NX-OS header lines that change on their own, !Time: and !Running configuration last done at:, so that the backups’ history only records real changes. This third one wasn’t in the filter. It never affected the classification, because it isn’t an owned line. But it put a commit into the audit trail that said “changed” when nothing had.

The fix went through the normal process, a pull request in iac-model with CI and a review:

volatile_headers: '(?m)^!((Time|Running configuration last done at):|No configuration change since last restart).*\n'

The colon stays required for the first two, so an unrelated comment that happens to start with !Time isn’t removed. NX-OS dropped the line itself at the next config change: the full deployment’s backups commit, c91bf36, removes it along with adding the NTP server. So the filter hasn’t had to remove it yet. It will after the next reload.

Notification#

Gitea can mail the user who started a failed run. In this lab it first didn’t, for two reasons:

  • Mail was disabled (Mailer Enabled: false). In the lab, Gitea now sends to a mail catcher (Mailpit) running next to it, which keeps every mail and delivers none, so nothing leaves the lab network.
  • Actions mails need their own switch. Mail for workflow runs arrived in Gitea 1.25 (go-gitea/gitea#34982), and in this lab it only sent with ENABLE_NOTIFY_MAIL = true in the [service] section. A working test mail from the admin page didn’t prove that part.

With both in place, the two failed drift runs I had started by hand each produced one mail, Run failed: drift.yml (11352e5317). One was the Cumulus01 outage, one was drift. The subject is the same for both, and the commit in it is the iac-deploy workflow’s, not the model’s. Green and pending runs, and the deployments, sent nothing.

The scheduled run that caught the hand change at 15:00 sent no mail at all. The discussion on that Gitea change notes that a scheduled run has no real user to send it to. So in this setup, the job nobody watches records drift and turns red, but nobody is told. A real setup needs its own alert, for example a final workflow step that calls the team’s chat webhook when the job fails. The lab has no chat to send to, so this one stays open.

Testing the boundary#

Two more tests check the boundary itself.

A branch can’t use the deploy runner. I pushed a branch to iac-model with this workflow, which is exactly the attack the design is meant to stop:

name: steal
on: push
jobs:
  grab:
    runs-on: deploy
    steps:
      - run: cat /etc/iac/vault_pass

The same push started the normal CI workflow, which the organization runner picked up and finished in 1 minute 46 seconds. The steal run stayed at Waiting, with a total duration of 0 seconds, for over an hour, until I cancelled it. No runner reachable from iac-model has the deploy label, so Gitea never handed the job out. Nothing about this depends on someone noticing the workflow in review.

The second engineer can’t deploy. Signed in as reviewer01, the deploy and drift workflows show their full run history, but neither shows a Run workflow button. The team can see every deployment and its log. Starting one is reserved for deployers.

The workflows#

Both files, exactly as they are on iac-deploy main after the fix:

name: deploy
on:
  workflow_dispatch:
    inputs:
      limit:
        description: "Device to deploy to, or all"
        required: true
        default: Arista01
        type: choice
        options: [Arista01, NXOS, Cumulus01, all]
      allow_removals:
        description: "Allow the plan to remove lines from devices"
        required: true
        default: false
        type: boolean

# One collection or deployment against the lab at a time, across deploy and drift.
concurrency:
  group: lab-devices
  cancel-in-progress: false

env:
  # Device credentials: a file only act-deploy can read. Never a Gitea secret.
  ANSIBLE_VAULT_PASSWORD_FILE: /etc/iac/vault_pass
  IAC_BACKUP_DIR: ${{ github.workspace }}/backups
  MODEL_SSH: ssh -i /var/lib/act-deploy/.ssh/model_ro -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222
  BACKUPS_SSH: ssh -i /var/lib/act-deploy/.ssh/backups_rw -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222
  LIMIT: ${{ github.event.inputs.limit }}
  ALLOW_REMOVALS: ${{ github.event.inputs.allow_removals }}

jobs:
  deploy:
    runs-on: deploy
    steps:
      - name: Model at main (reviewed code only), and the backups
        run: |
          rm -rf model backups
          GIT_SSH_COMMAND="$MODEL_SSH" git clone -q --depth 1 --branch main [email protected]:netops/iac-model.git model
          GIT_SSH_COMMAND="$BACKUPS_SSH" git clone -q --branch main [email protected]:netops/iac-backups.git backups
          git -C backups config user.name "iac-pipeline"
          git -C backups config user.email "[email protected]"
          echo "model main: $(git -C model rev-parse --short HEAD)  limit: $LIMIT  allow_removals: $ALLOW_REMOVALS"

      - name: Toolchain (locked)
        working-directory: model
        run: uv sync --locked && mkdir -p logs

      - name: Pre-deploy collection
        working-directory: model
        run: |
          uv run --locked ansible-playbook playbooks/backup.yml --limit "$LIMIT"
          sha=$(git rev-parse --short HEAD)
          git -C "$IAC_BACKUP_DIR" add -A
          git -C "$IAC_BACKUP_DIR" diff --cached --quiet && echo "devices unchanged since the last collection" || {
            git -C "$IAC_BACKUP_DIR" commit -q -m "pre-deploy $sha ($LIMIT)"
            GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
            echo "devices had changed since the last collection - committed as pre-deploy $sha"; }

      - name: Plan and guardrails
        id: plan
        working-directory: model
        run: |
          uv run --locked ansible-playbook playbooks/diff.yml --limit "$LIMIT"
          for f in remediation/*.cfg; do
            if [ -s "$f" ]; then echo "== $(basename "$f" .cfg)"; cat "$f"; else echo "== $(basename "$f" .cfg): nothing to send"; fi
          done
          # One summary per selected device, or the step fails. No summary must never read as "nothing to do".
          hosts=$(uv run --locked ansible eos,nxos,cumulus --list-hosts --limit "$LIMIT")
          hosts=$(tail -n +2 <<< "$hosts")
          test -n "$hosts" || { echo "no devices match '$LIMIT'"; exit 1; }
          counts=$(python3 -c 'import json,sys
          m=e=0
          for h in sys.argv[1:]:
              d=json.load(open("remediation/%s.json" % h)); m+=int(d["missing"]); e+=int(d["extra"])
          print(m, e)' $hosts)
          read missing extra <<< "$counts"
          echo "plan: missing=$missing extra=$extra"
          if [ "$missing" = 0 ] && [ "$extra" = 0 ]; then echo "already converged - nothing to send"; echo "converged=true" >> "$GITHUB_OUTPUT"; exit 0; fi
          if [ "$extra" != 0 ] && [ "$ALLOW_REMOVALS" != "true" ]; then
            echo "REFUSED: the plan removes $extra line(s). A removal can mean someone added something by hand."
            echo "Check the lines above. If removing them is intended, run again with allow_removals."
            exit 1
          fi
          echo "converged=false" >> "$GITHUB_OUTPUT"

      - name: Apply, re-collect and verify (the convergence gate)
        id: apply
        if: steps.plan.outputs.converged == 'false'
        working-directory: model
        run: uv run --locked ansible-playbook playbooks/apply.yml --limit "$LIMIT" -e pipeline_confirmed=true

      # Runs whenever the plan reached a decision, including after a failed apply.
      # DEPLOYED: the last full deployment that sent something and passed the gate.
      # CONVERGED: the last model all devices were verified against (gate, no-op deploy, or a green drift check).
      - name: Record the result in the backups
        if: always() && steps.plan.outputs.converged != ''
        working-directory: model
        run: |
          sha=$(git rev-parse --short HEAD)
          stamp() { printf 'model %s\n%s %s\nrun %s/%s/actions/runs/%s\n' "$(git rev-parse HEAD)" "$1" "$(date -u +%FT%TZ)" "${{ github.server_url }}" "${{ github.repository }}" "${{ github.run_number }}"; }
          last=$(sed -n 's/^model //p' "$IAC_BACKUP_DIR/CONVERGED" 2>/dev/null || true)
          if [ "${{ steps.plan.outputs.converged }}" = true ]; then
            msg="verified $sha ($LIMIT, already converged)"
            if [ "$LIMIT" = all ] && [ "$last" != "$(git rev-parse HEAD)" ]; then stamp verified > "$IAC_BACKUP_DIR/CONVERGED"; fi
          else
            if [ "${{ steps.apply.outcome }}" != success ]; then
              echo "apply did not succeed - collecting again, so the record shows the state after the attempt"
              uv run --locked ansible-playbook playbooks/backup.yml --limit "$LIMIT" || echo "post-attempt collection incomplete: an unreachable device keeps its pre-deploy file"
            fi
            msg="post-deploy $sha ($LIMIT, gate: ${{ steps.apply.outcome }})"
            if [ "$LIMIT" = all ] && [ "${{ steps.apply.outcome }}" = success ]; then
              stamp deployed > "$IAC_BACKUP_DIR/DEPLOYED"
              stamp verified > "$IAC_BACKUP_DIR/CONVERGED"
            fi
          fi
          git -C "$IAC_BACKUP_DIR" add -A
          git -C "$IAC_BACKUP_DIR" diff --cached --quiet && { echo "nothing to record"; exit 0; }
          git -C "$IAC_BACKUP_DIR" commit -q -m "$msg"
          GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
          git --no-pager -C "$IAC_BACKUP_DIR" log --oneline -1
name: drift
on:
  schedule:
    - cron: "*/30 * * * *"   # Gitea runs schedules in UTC; every 30 min doesn't care
  workflow_dispatch:

concurrency:
  group: lab-devices
  cancel-in-progress: false

env:
  ANSIBLE_VAULT_PASSWORD_FILE: /etc/iac/vault_pass
  IAC_BACKUP_DIR: ${{ github.workspace }}/backups
  MODEL_SSH: ssh -i /var/lib/act-deploy/.ssh/model_ro -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222
  BACKUPS_SSH: ssh -i /var/lib/act-deploy/.ssh/backups_rw -o IdentitiesOnly=yes -o UserKnownHostsFile=/var/lib/act-deploy/.ssh/known_hosts -p 2222

jobs:
  drift:
    runs-on: deploy
    steps:
      - name: Model at main, and the backups
        run: |
          rm -rf model backups
          GIT_SSH_COMMAND="$MODEL_SSH" git clone -q --depth 1 --branch main [email protected]:netops/iac-model.git model
          GIT_SSH_COMMAND="$BACKUPS_SSH" git clone -q --branch main [email protected]:netops/iac-backups.git backups
          git -C backups config user.name "iac-pipeline"
          git -C backups config user.email "[email protected]"

      - name: Toolchain (locked)
        working-directory: model
        run: uv sync --locked && mkdir -p logs

      - name: Collect, and commit only what changed
        working-directory: model
        run: |
          uv run --locked ansible-playbook playbooks/backup.yml
          git -C "$IAC_BACKUP_DIR" add -A
          if git -C "$IAC_BACKUP_DIR" diff --cached --quiet; then echo "no device changed since the last collection"; else
            changed=$(git -C "$IAC_BACKUP_DIR" diff --cached --name-only | tr '\n' ' ')
            git -C "$IAC_BACKUP_DIR" commit -q -m "drift-check: changed on $changed"
            GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
            echo "committed: changed on $changed"
          fi

      - name: Compare with the model, and classify
        working-directory: model
        run: |
          uv run --locked ansible-playbook playbooks/diff.yml >/dev/null
          # One summary per device, or the step fails. No summary must never read as green.
          hosts=$(uv run --locked ansible eos,nxos,cumulus --list-hosts)
          hosts=$(tail -n +2 <<< "$hosts")
          test -n "$hosts" || { echo "no devices in the inventory"; exit 1; }
          counts=$(python3 -c 'import json,sys
          m=e=0
          for h in sys.argv[1:]:
              d=json.load(open("remediation/%s.json" % h)); m+=int(d["missing"]); e+=int(d["extra"])
          print(m, e)' $hosts)
          read missing extra <<< "$counts"
          model=$(git rev-parse HEAD)
          # The last model all devices were verified against. Before the first CONVERGED file, fall back to DEPLOYED.
          verified=$(sed -n 's/^model //p' "$IAC_BACKUP_DIR/CONVERGED" 2>/dev/null || true)
          [ -n "$verified" ] || verified=$(sed -n 's/^model //p' "$IAC_BACKUP_DIR/DEPLOYED" 2>/dev/null || true)
          for f in remediation/*.cfg; do if [ -s "$f" ]; then echo "== $(basename "$f" .cfg)"; cat "$f"; fi; done
          if [ "$missing" = 0 ] && [ "$extra" = 0 ]; then
            echo "GREEN: owned lines match model ${model:0:7} on all devices"
            if [ "$model" != "$verified" ]; then
              printf 'model %s\nverified %s\nrun %s/%s/actions/runs/%s\n' "$model" "$(date -u +%FT%TZ)" "${{ github.server_url }}" "${{ github.repository }}" "${{ github.run_number }}" > "$IAC_BACKUP_DIR/CONVERGED"
              git -C "$IAC_BACKUP_DIR" add CONVERGED
              git -C "$IAC_BACKUP_DIR" commit -q -m "verified ${model:0:7} (drift-check, all devices)"
              GIT_SSH_COMMAND="$BACKUPS_SSH" git -C "$IAC_BACKUP_DIR" push -q origin main
              echo "recorded: CONVERGED is now ${model:0:7}"
            fi
            exit 0
          fi
          if [ "$model" != "$verified" ]; then
            echo "PENDING DEPLOYMENT: model main ${model:0:7} differs from the last verified model (${verified:-none recorded yet})."
            echo "Difference: missing=$missing extra=$extra. Not an incident - deploy when ready."
            exit 0
          fi
          echo "DRIFT: devices differ from model ${verified:0:7}, which they were last verified against: missing=$missing extra=$extra."
          echo "Someone changed a device outside the pipeline. The backups commit above shows what changed."
          exit 1

The audit trail#

The backups’ history for the day, all written by the pipeline:

8be96eb 15:28 post-deploy 0fc0e79 (Arista01, gate: success)
781a30e 15:01 drift-check: changed on Arista01
d7e35ea 14:44 verified 0fc0e79 (drift-check, all devices)
d6acc0a 14:42 post-deploy 0fc0e79 (Arista01, gate: success)
8d101b4 14:31 drift-check: changed on Arista01
23fe862 09:10 post-deploy 6b90a7a (all, gate: success)
28855c2 08:52 post-deploy ca8cbf4 (Arista01, gate: success)
deb0b67 08:27 drift-check: changed on Arista01
c91bf36 08:06 post-deploy 85c2c99 (all, gate: success)
661f5d4 06:39 post-deploy 85c2c99 (Arista01, gate: success)
9b9aa76 06:30 drift-check: changed on Arista01
e7c0cce 06:25 post-deploy 85c2c99 (Arista01, gate: success)
e79e918 06:10 drift-check: changed on NXOS

Read from the bottom. Morning, first version: a reboot that looked like a change, the first deployment to one device, a hand change and its removal, the full NTP deployment, the hand change again (red this time), its removal, and the syslog change deployed to all three. Afternoon: the hand change during the gap (collected at 14:30 by the first version, minutes before the fix was merged), its removal, the first verification of 0fc0e79, the hand change caught by the schedule, and its removal.

What this does not show#

  • One owner. The deploy workflows, the branch rules and the runner registration are all controlled by one person, and the workflow fix was merged without a second reviewer. The design needs at least two owners before it means anything in a team.
  • The deploy runner can read the key. That’s its job. What protects the key is that only owner-controlled workflows run on that runner, and they only check out main. Anyone with root on the runner host can read it too, so that host is part of the trust boundary.
  • The runner isn’t the only copy of the key. The key holder still has one, for running playbooks by hand. It’s the only copy a CI runner can reach.
  • The model is trusted once it’s on main. A reviewed pull request can change a playbook, and the deploy runner will run it with the key. The review in part 2 is the control. It’s a person, not a machine.
  • Only owned lines are checked. “Converged” means the NTP and syslog server lines match the model. It says nothing about the rest of the device config, and nothing about whether NTP is synchronized or syslog messages arrive.
  • Pending can hide drift. While a merged change waits for deployment, or before the first green run after any model change, a hand change shows as 🟡, because the job can’t tell which difference came from where. It stays that way until a run verifies the devices against the new model. The next deployment does surface it: the plan lists the extra line and refuses to remove it without allow_removals.
  • Scheduled failures alert nobody. See the notification section above.
  • Two fixes tested offline only: one summary per device, and collecting again after a failed apply.
  • The deployers permissions are only half tested. reviewer01 has no Run workflow button, as intended. But the only deployer is darren, who is also an owner and would see the button anyway. A deployer who isn’t an owner hasn’t been tested.
  • Deployment is manual. Merging doesn’t deploy. That’s a choice for a lab with one owner, not a limitation of the tools.
  • Trust on first use for the Gitea SSH host key, which the deploy runner learned with ssh-keyscan on the lab network. And plain HTTP between runner and Gitea, as in part 2.

Conclusion#

Both approved changes are on all three devices, and their owned lines match the model at main. For the lines this part deployed, the backups lead from the device back to the deployment run that sent them, the model commit, and the pull request that approved it. The drift job tells a waiting change from a hand change once it has verified the current model. The first version of it didn’t, and a review of this write-up caught that before the post went out.


Next in this series: so far the pipeline manages two services on three devices. The next parts take it to other platforms and other IaC scenarios.

← more in Automation