darrenqu.net

Automation

Network IaC, Part 2: CI a Team Can Trust

4338 words 21 min read

ansibleci

Putting the Ansible model from part 1 under Gitea Actions, twice: once for one person, then again for a team. What a clean checkout found, where the secrets were leaking, and how to compute a full change plan without ever handing CI a device password.

on this page

Network IaC from a brownfield lab. Part 2 of a series. Part 1 built the Ansible workflow: collect, model, diff offline, apply, re-collect. This part puts it under CI, for a team. Every branch is checked in a clean workspace, every change shows the exact commands it would send, and nothing reaches main without a second engineer’s approval. Part 3 covers deployment and scheduled drift detection.

Part 1 ended on a list of bugs that Ansible accepted without complaint: four misspelled setting names and one invalid value. The obvious fix is to have a machine read the files on every change.

I built that twice. The first version worked, and it was built for one person: my token, my runner user, my repositories. When I asked what would happen if a second engineer joined, most of it fell apart. The second version is built for a team, and I’m one user in it. This post covers both, because most of what I learned came from the gap between them.

Lab setup#

  • Same lab as part 1: Arista vEOS 4.32.4M, Cisco Nexus 9300v 10.6.3, Cumulus VX 5.9.1 in EVE-NG, one device each, NTP and syslog only.
  • CI platform: Gitea on my own network, 1.22.6 at the start and upgraded to 1.27.1 during this part, with Gitea Actions and one runner (act_runner 0.2.11, which kept working across the upgrade and reconnected on its own). The workflow is GitHub Actions syntax, with nothing Gitea-specific in the file.
  • The “team” is partly simulated. There’s a second engineer account (reviewer01) and a service account (iac-bot). I control all three, so the review requirement is demonstrated here, not independently exercised.
  • This part is CI only. Nothing here writes to a device.

The first version, and what was wrong with it#

The first version ran on every push: syntax check, then a “plan” that diffed the model against the committed backups. It found real bugs, covered below. But read as a team setup, it had four problems, and none of them showed up as an error:

  1. Every job could read the vault key. The runner ran as ubuntu on the lab’s control node, the same user that kept ~/.vault_pass for running playbooks by hand. In host mode, a job sees whatever its Unix user sees, so any workflow on any branch could have printed the key.
  2. CI depended on my personal token. It read the backups repository with an access token I’d created on my own account. If I left, CI would stop.
  3. Every branch push got the vault key as a secret. A branch isn’t reviewed. In a team, anyone who can push a branch can write a workflow that prints its secrets.
  4. Both repositories were public. I’d recreated them and missed the private checkbox. That includes the backups repo, which holds password hashes. I didn’t notice until a check passed when it shouldn’t have.

The fourth is worth a closer look, because the check designed to catch a broken token is what hid it. The token check asked the API whether the token could read the backups repo, and got 200, while the token was in fact empty. With no token, curl asked anonymously, and a public repository answers anyone. The check now fails if the token is empty, before it asks anything:

test -n "$T" || { echo "BACKUPS_TOKEN is empty: org secret missing or misnamed"; exit 1; }

And the repositories’ privacy is now its own check, run from outside without logging in: a private repository answers 404.

The team design#

Code stateWho can trigger itWhat it gets
any branch (not reviewed)any engineerread access to the backups. Never device credentials
main (reviewed, merged)only an approved pull requestdevice credentials, for deployment (part 3)

A trigger alone can’t enforce the second row. Gitea secrets aren’t tied to a branch, and anyone who can push a branch can write their own workflow file. So in part 3, the vault key is never a Gitea secret at all. It’s a file on the runner host, readable only by a separate runner that is registered to a separate deployment repository. That runner checks out the model at main, and only owners can change its workflows. This part builds everything up to that boundary.

The rest follows from that:

Setup
Organizationnetops, private. Repositories belong to the team that’s responsible for them, not to a person or a tool
Peopleeach has their own account. darren is an owner and an engineer; reviewer01 is an engineer
Service accountiac-bot, marked restricted, so it sees only what a team explicitly gives it. Used interactively only to create or rotate its token
Teamsengineers (write iac-model), readers (read both repos), ci-bots (read iac-backups only)
Runnerregistered to the organization, running as a dedicated system user act-runner, which holds no device credentials
Workstationsedit, check, push branches. No vault key, no route to the devices

Two details are easy to get wrong. A Gitea team has one permission level for all its repositories, so “engineers write the model and read the backups” takes two teams, engineers and readers. And a token for a restricted account set to “all repositories” can still only reach what that account’s teams allow:

iac-backups: 200
iac-model:   404

The bot can read exactly one repository.

Getting device passwords out of CI’s reach#

The plan job diffs the model against the backups. It never logs in to a device. So in principle it shouldn’t need the vault key at all. In practice it did, for two separate reasons.

Reason 1: ansible.cfg pointed at a key file. With vault_password_file = ~/.vault_pass in the shared config, every Ansible command fails on a machine without that file, even a syntax check:

[ERROR]: The vault password file /home/…/.vault_pass was not found

So a workstation without the key couldn’t run any check, and my first CI’s “no secrets” job only passed because the runner happened to be the machine that had the file. The key path now comes from ANSIBLE_VAULT_PASSWORD_FILE, set only where decryption is legitimate.

Reason 2: the password pointer in group_vars. The device password was wired up the conventional way: an encrypted vault.yml in each platform’s group_vars, and a readable pointer in group_vars/all:

ansible_password: "{{ vault_ansible_password }}"

The obvious fix was to move the vault files somewhere Ansible doesn’t load automatically. I tested that in a throwaway copy. The run got through the first few local tasks, then failed partway through the offline diff, on tasks that are delegated to the control node and never connect to a device:

Error processing keyword 'password': 'vault_ansible_password' is undefined

So Ansible resolved the connection’s password for tasks that weren’t going to connect. I didn’t trace exactly which tasks trigger that and which don’t; what matters here is that, as long as the pointer exists, the diff can fail without its target. So the pointer had to go. The encrypted file now sets ansible_password itself, and lives in secrets/<platform>.yml, which Ansible never loads on its own. Only the playbook that logs in loads it:

  tasks:
    - name: Refuse to run without the vault key
      ansible.builtin.assert:
        that: lookup('env', 'ANSIBLE_VAULT_PASSWORD_FILE') | length > 0
        fail_msg: "This playbook logs in to devices. Set ANSIBLE_VAULT_PASSWORD_FILE (runner and key holder only)."

    - name: Load device credentials (only playbooks that log in do this)
      ansible.builtin.include_vars:
        file: "{{ playbook_dir }}/../secrets/{{ platform }}.yml"

Variables from include_vars stay set for the rest of the run, so apply.yml, which starts by importing backup.yml, has the credentials in its later plays too. I checked that on the devices before relying on it. Without the assert, a missing key failed with Unknown error, which is not a message anyone should have to debug.

Moving the passwords meant rewriting encrypted files, which only the key holder can do. The plaintext only ever passed through a pipe. set -euo pipefail makes a failure anywhere in the pipe stop the script, and the original is only removed after the new file has been checked: it must decrypt, and parse as YAML with exactly one key, the renamed one, holding a non-empty string:

set -euo pipefail
ansible-vault view inventory/group_vars/eos/vault.yml \
  | sed 's/^vault_ansible_password:/ansible_password:/' \
  | ansible-vault encrypt --output secrets/eos.yml -
ansible-vault view secrets/eos.yml | python3 -c '
import sys, yaml
d = yaml.safe_load(sys.stdin)
ok = isinstance(d, dict) and list(d) == ["ansible_password"] \
     and isinstance(d["ansible_password"], str) and d["ansible_password"] != ""
sys.exit(0 if ok else "secrets/eos.yml: expected exactly one non-empty ansible_password")'
git rm -q inventory/group_vars/eos/vault.yml

That’s the version I’d use now. When I did the migration, the check was a manual look at the decrypted key names, with the values masked.

After the move, the backups collected with the key were byte-identical to before, and the full diff ran with no key at all:

owned=2 missing=0 extra=0     (all three devices, no vault key on the machine)

The inventory checker enforces the new rule, within what it actually reads: inventory/group_vars/<group>/*.yml and the host entries in the inventory file. In those, a credential key (ansible_password, or anything starting with vault_) fails CI, and so does a vault.yml. It also requires secrets/<platform>.yml to exist for each platform and to start with an Ansible Vault header. It doesn’t read host_vars/, .yaml files or nested directories. None exist in this repository yet; widening the checker is the job when they do.

Why CI runs on pushes, not pull requests#

The usual design is to run checks on the pull_request event. In my first attempt the plan job failed there, and the log showed why:

triggered by event: pull_request
expression 'format('{0}', secrets.BACKUPS_TOKEN)' evaluated to '%!t(string=)'

The secret was empty in my AGit PR run, so I switched to branch pushes for the plan job.

So CI runs on pushes to any branch. Pushing a branch requires write access, so the person is trusted to that extent. The pull request is then opened from the branch, and Gitea shows the branch’s CI result on it, because it’s the same commit. I’d first opened pull requests with Gitea’s AGit flow, which creates a pull request without a branch. That never triggers a push run, so it doesn’t fit this design.

What each job gets, as a result:

JobRuns onSecretsPurpose
checksevery branchnonesyntax, lint, inventory and config checks
planevery branchthe bot’s read-only backups tokenthe exact commands this commit would send
deploy (part 3)main onlydevice credentialsapply, re-collect, verify

Engineers can read the bot’s token by writing a workflow on a branch. That gets them read access to the backups, which they already have through readers. It grants nothing new, and it’s why the token is scoped that narrowly.

plan checks itself as well. My first version of that check only tested one environment variable, and named the step “No device credentials in this job”, which claimed far more than it proved. A key could still come from a config file or a vault ID. The current version keeps the variable check under an honest name, and adds a check that tries to decrypt one of the credential files and requires the specific failure:

      - name: Device credentials cannot be decrypted here
        run: |
          # ansible (not ansible-vault) never prompts for a vault password, so this can't hang.
          # here-string, not 'echo | grep -q': under pipefail, grep -q exiting early can SIGPIPE echo.
          # An empty secrets/ fails too: the unmatched glob makes ansible fail, so it's 'inconclusive'. n is a second guard.
          n=0
          for f in secrets/*.yml; do
            out=$(timeout 60 uv run --locked ansible localhost -m ansible.builtin.include_vars -a "file=$f" -vvv </dev/null 2>&1) && { echo "plan can decrypt $f"; exit 1; }
            grep -qF "no vault secrets found" <<< "$out" || { echo "check inconclusive for $f:"; echo "$out" | tail -5; exit 1; }
            n=$((n+1))
          done
          test "$n" -gt 0 || { echo "no files in secrets/ - nothing was checked"; exit 1; }
          echo "ok: none of the $n files in secrets/ can be decrypted in this job"

It tries every .yml file in secrets/. If any of them decrypts, the job fails. Any result other than the expected “no vault secrets found” fails it as inconclusive, so a broken check can’t look like a passing one. The comments in the step explain the shell details. In CI it reports ok: none of the 3 files in secrets/ can be decrypted in this job. I ran the same command as the key holder first, to see it succeed there.

The first attempt at this used ansible-vault view, and it hung. Without a key, ansible-vault asks for the password interactively. There’s no keyboard in a CI job, so the step didn’t fail, it waited: four minutes before I found the stuck process on the runner. ansible itself never prompts unless it’s told to, and the timeout makes sure that whatever goes wrong next, the step ends.

Getting this change in took two pull requests. The first one’s commit message described the whole change, but the commit itself only contained a one-line comment edit: I’d run the second half of the patch and not the first. The review approved the description, and the diff would have shown the gap. The second pull request applied the change for real, and its message says so. Branch protection guarantees that someone approves a change; it doesn’t guarantee that they read the diff.

A runner that belongs to the team#

The runner is the program that actually executes jobs. Gitea stores repositories and secrets and queues jobs, and runs nothing itself. act_runner polls Gitea for work, so the connection goes from the lab outward, never into it.

In the team version it runs as a dedicated system user:

  • act-runner, with its home in /var/lib/act-runner, and the binary and config in system paths
  • registered to the organization, not to a person
  • unable to read the key holder’s home directory
$ sudo -u act-runner cat /home/ubuntu/.vault_pass >/dev/null 2>&1 \
    && echo "PROBLEM: act-runner can read the vault key" \
    || echo "ok: act-runner cannot read /home/ubuntu"
ok: act-runner cannot read /home/ubuntu
$ ls -ld /home/ubuntu; id act-runner
drwxr-x--- 9 ubuntu ubuntu 4096 Oct  8 09:18 /home/ubuntu
uid=999(act-runner) gid=988(act-runner) groups=988(act-runner)

The runner’s own credentials (.runner) are mode 0600, so other users on the host can’t read them. That doesn’t protect them from the jobs themselves: jobs run as act-runner, the file’s owner. A workflow could read .runner, and with it register as this runner and receive other jobs. It could also leave processes or files behind for the next job. That’s acceptable here only because the runner holds no device credentials and the people who can push are trusted.

Its name (iac-runner01) identifies the machine. Its label (lab) says what it can do: reach the lab network. Workflows ask for the capability with runs-on: lab. A second runner with the same label would share the work without any workflow changing.

Setting it up took three tries, each with a slightly misleading error. A registration token is invalidated by the next click on “Create new runner”. --labels on the command line is silently overridden by the labels in the generated config.yaml. And those generated labels are Docker images, so a runner with no Docker exits at startup and systemd retries until it gives up.

What a clean checkout found#

Independent of the team design, running the playbooks in an empty directory found three problems that had never failed on the control node:

  • A playbook that didn’t parse. An edit swapped two indentation levels in backup.yml. My local check after the edit passed, because it ran a different playbook. CI syntax-checks every playbook, so it found the error on the first run.
  • Output directories that only existed on one machine. rendered/ and remediation/ are git-ignored, and the playbooks assumed they existed. They now create them.
  • A pager and a log path. git log waited for input (Press RETURN), and Ansible couldn’t write logs/ansible.log.

These checks found assumptions I’d missed because I always ran the playbooks on the same control node.

Checks that fail on what Ansible accepts#

  • A locked toolchain. Every uv sync and uv run in CI uses --locked, so a job fails if uv.lock doesn’t match pyproject.toml. I’d started with uv sync --frozen followed by plain uv run: --frozen could miss a mismatch with pyproject.toml, and a later plain uv run could update the lockfile during CI.
  • ansible-lint found eight problems on its first run, and all of them were worth fixing. The playbooks now pass the strictest profile, production.
  • An allowlist for inventory keys. Every key in group_vars and the inventory must be a known Ansible connection variable or a known model variable. A KNOWN_VRFS table catches the value typo (defatult), which no tool could know is wrong. The same checker enforces where credentials may live.
  • The effective config, not the key names. ansible-config validate reported my correct callback_result_format as an unknown key. It only knows core settings, not plugin options. So the check asserts the effective value with ansible-config dump -t callback --only-changed.

Each check was tested by putting its bug back, one at a time:

== ansible_become_mothod: enable               FAIL unknown key 'ansible_become_mothod'
== nsibel_password: x                          FAIL unknown key 'nsibel_password'
== mgmt_vrf: defatult                          FAIL mgmt_vrf 'defatult' is not a known VRF for eos
== ansibel_host                                FAIL unknown key 'ansibel_host' / no ansible_host
== host named after its group                  FAIL host 'cumulus' has the same name as a group
== ansible_password pointer back in all/vars   FAIL 'ansible_password' is a credential …
== vault.yml back in group_vars                FAIL no vault files in group_vars …

The first time I ran this, one injection silently did nothing: the sed expected the wrong indentation, so the bug never went in. That looked exactly like a check passing. The loop now confirms each injection changed a file before running the checker.

Review that holds, and where it doesn’t#

main on both repositories is protected:

  • No direct pushes, owners included. A test push from my owner account was refused by the server:
    remote: Gitea: Not allowed to push to protected branch main
     ! [remote rejected] main -> main (pre-receive hook declined)
  • One approval required, only counted from the engineers team, and dismissed if the branch changes afterwards. Every pull request merged in this part was approved by reviewer01, never by its author. I didn’t separately test whether Gitea refuses an author’s approval of their own pull request.
  • Both CI jobs must be green. I came close to locking every merge out here: the list of known status checks still showed ci / checks (pull_request) from before the trigger change. Requiring that context would have waited forever for a run that no longer happens. Only (push) is required.

The first real change went through the whole path. I pushed the branch demo/ntp-011 as darren, CI planned it, reviewer01 approved it in a separate session, and it was merged. Its plan, computed with no device password anywhere in the job:

backups snapshot a8495a6 from 2026-09-30T07:12:08+00:00 - NXOS: typo admin account darrem removed
owned=2 missing=1 extra=0     (each device)
== Arista01
ntp server 192.168.100.11
== Cumulus01
nv set service ntp mgmt server 192.168.100.11
nv config apply -y
nv config save
== NXOS
ntp server 192.168.100.11 use-vrf management

The second branch reintroduced the become_mothod typo. Its checks job failed at the inventory check, and plan was skipped because it depends on checks. On the pull request, after the upgrade described below, Gitea listed every reason it couldn’t be merged:

✕ ci / checks (push)   Failing after 41s        Required
⊘ ci / plan (push)     Has been skipped         Required
✕ Some required checks were not successful.
✕ This pull request doesn't have enough required approvals yet. 0 of 1 approvals granted from users or teams on the allowlist.
✕ This pull request is blocked because it's outdated.

I closed the pull request without merging it.

Where it didn’t hold, and the upgrade. On the pull requests I opened, Gitea 1.22.6 showed me, as an owner:

This pull request doesn't have enough approvals yet. 0 of 1 approvals granted.
As an administrator, you may still merge this pull request.

with an active merge button. So the review requirement bound engineers, not owners. Gitea added a branch protection setting to block that override in 1.23. My server was several releases behind, and those releases also carried security fixes relevant to this setup: one let organization secrets be extracted through the fork API, another let a fork’s Actions job read a third private repository. So I upgraded to 1.27.1. I backed up first (a database dump plus the data directories), then pinned the exact image tag.

After the upgrade, the protection rule as the API reports it:

netops/iac-model   main
  enable_push                  False
  required_approvals           1
  approvals_whitelist_teams    ['engineers']
  dismiss_stale_approvals      True
  block_on_rejected_reviews    True
  block_on_outdated_branch     True
  status_check_contexts        ['ci / checks (push)', 'ci / plan (push)']
  block_admin_merge_override   True

On the typo pull request above, which had no approval and failing checks, the line “As an administrator, you may still merge” no longer appeared for my owner account. The only merge option Gitea offered was to schedule a merge for when checks succeed. I didn’t test whether that scheduled merge also waits for the approval, so I didn’t click it.

The next pull request, a README describing this workflow, went through the same path as everyone’s: two commits with green CI, an approval and a comment from reviewer01, then the merge. An organizational rule is still worth keeping alongside the setting: administer with one account and work with another, and have at least two owners.

The workflow#

name: ci
on:
  push:
    branches: ['**']

# Branch pushes let plan use the bot's read-only backups token. It never gets device credentials.
jobs:
  # Static checks: no secrets, no backups, no devices. Fails fast on typos.
  checks:
    runs-on: lab
    steps:
      - name: Checkout model
        uses: actions/checkout@v4

      - name: Toolchain (locked)
        run: uv sync --locked

      - name: Local log directory
        run: mkdir -p logs

      - name: Syntax check every playbook
        run: for p in playbooks/*.yml; do uv run --locked ansible-playbook "$p" --syntax-check || exit 1; done

      - name: Lint
        run: uv run --locked ansible-lint playbooks/

      - name: Inventory keys, values and credential placement
        run: uv run --locked python tests/check_inventory.py

      - name: Effective config
        run: uv run --locked tests/check_config.sh

  # Plan: what applying this commit would send, against the last collected backups.
  # Drift here is expected for a real change, so it is reported, not failed on.
  plan:
    needs: checks
    runs-on: lab
    env:
      IAC_BACKUP_DIR: ${{ github.workspace }}/_backups
    steps:
      - name: No vault key in the environment
        run: |
          test -z "${ANSIBLE_VAULT_PASSWORD_FILE:-}" || { echo "ANSIBLE_VAULT_PASSWORD_FILE is set in plan"; exit 1; }
          echo "ok: ANSIBLE_VAULT_PASSWORD_FILE not set"

      - name: Checkout model
        uses: actions/checkout@v4

      - name: Check BACKUPS_TOKEN can read the backups repo
        env:
          T: ${{ secrets.BACKUPS_TOKEN }}
        run: |
          test -n "$T" || { echo "BACKUPS_TOKEN is empty: org secret missing or misnamed"; exit 1; }
          code=$(curl -s -o /dev/null -w "%{http_code}" -H "Authorization: token $T" http://10.0.20.22:3000/api/v1/repos/netops/iac-backups)
          echo "backups API: $code"
          test "$code" = 200

      - name: Checkout backups (the snapshot this plan is computed against)
        uses: actions/checkout@v4
        with:
          repository: netops/iac-backups
          ref: main
          token: ${{ secrets.BACKUPS_TOKEN }}
          path: _backups

      - name: Record which snapshot
        run: git --no-pager -C _backups log -1 --format="backups snapshot %h from %cI - %s"

      - name: Local log directory
        run: mkdir -p logs

      - name: Toolchain (locked)
        run: uv sync --locked

      - name: Device credentials cannot be decrypted here
        run: |
          # ansible (not ansible-vault) never prompts for a vault password, so this can't hang.
          # here-string, not 'echo | grep -q': under pipefail, grep -q exiting early can SIGPIPE echo.
          # An empty secrets/ fails too: the unmatched glob makes ansible fail, so it's 'inconclusive'. n is a second guard.
          n=0
          for f in secrets/*.yml; do
            out=$(timeout 60 uv run --locked ansible localhost -m ansible.builtin.include_vars -a "file=$f" -vvv </dev/null 2>&1) && { echo "plan can decrypt $f"; exit 1; }
            grep -qF "no vault secrets found" <<< "$out" || { echo "check inconclusive for $f:"; echo "$out" | tail -5; exit 1; }
            n=$((n+1))
          done
          test "$n" -gt 0 || { echo "no files in secrets/ - nothing was checked"; exit 1; }
          echo "ok: none of the $n files in secrets/ can be decrypted in this job"

      - name: Plan
        run: uv run --locked ansible-playbook playbooks/diff.yml

      - name: Commands this commit would send
        run: |
          for f in remediation/*.cfg; do
            if [ -s "$f" ]; then echo "== $(basename "$f" .cfg)"; cat "$f"; else echo "== $(basename "$f" .cfg): nothing to send"; fi
          done

What this does not show#

  • The plan is only as fresh as the last collection. It’s computed against the committed backups. A green plan means the diff completed against that snapshot (a8495a6 here), whatever it reports: the NTP plan was green while reporting a missing line on every device. And the snapshot may not be what the devices hold right now. The scheduled collection in part 3 is what keeps that snapshot current.
  • Host-mode isolation. Jobs run directly on the runner host as act-runner. That user holds no device credentials, but it does hold the runner’s own Gitea credentials, and every job can read them. An engineer’s branch can also run arbitrary commands on a host inside the lab network and leave processes or files behind. Container mode would reduce that, because each job starts in a fresh container. It isn’t a complete boundary, though: how much it isolates depends on what the container can reach and what is mounted into it.
  • One owner, and a simulated second engineer.
  • Plain HTTP between the runner and Gitea inside the lab.

Conclusion#

The NTP change is now approved and merged. CI showed the commands for all three devices, but none have been applied. Part 3 starts with deploying that change and checking the devices afterwards.


Next in this series: Part 3 deploys the merged change through a separate deployment repository, whose runner is the only place the device credentials exist. It runs the convergence gate after each deployment, and a scheduled job that collects, commits and fails on drift.

← more in Automation