Adopting a Brownfield Network into Ansible: What the Lab Caught
Three vendors, one config domain, no rewrite. Collect, model, diff offline, remediate, and re-collect: what failed and where it showed up.
on this page
Network IaC from a brownfield lab. Part 1 of a series. This part covers the Ansible workflow itself: project setup, collection, modeling, diff and apply. Part 2 turns it into a CI/CD pipeline: a remote Git repository, checks on every change, and gated deployment to the devices. Later parts cover other platforms and IaC scenarios.
Most network IaC tutorials start from an empty switch. Mine never are. The network already exists, it already carries traffic, and nobody wrote it down. The question is not “how do I template a new device” but “how do I get a device someone else configured years ago under control of a model, without a rewrite and without breaking it.”
This is a lab run of that process on three vendors. The approach is deliberately small:
- Collect the running config, read-only.
- Model what should be true, from what the backups show — not from memory.
- Diff offline: compare the model with the backup, never with the live device.
- Remediate: generate the exact commands, read them, push them to one device.
- Re-collect and re-diff. A change counts as converged when a fresh collection shows no missing and no extra owned lines.
The failures below showed up in different places. A VRF typo and unstable ordering showed up in the generated remediation, before anything was pushed. NX-OS rewriting a line only showed up at step 5: the device accepted the command without an error, and the re-collected config disagreed with the model. The Cumulus behavior showed up during a manual cleanup, outside the domain the model manages.
Scope, stated plainly#
- Lab, not production. EVE-NG, virtual images: Arista vEOS 4.32.4M, Cisco Nexus 9300v (NX-OS 10.6.3), NVIDIA Cumulus VX 5.9.1. One device each.
- One config domain: NTP and syslog servers. Top-level lines, not nested under an interface, whose order doesn’t matter. That is the easiest domain there is, chosen so the pipeline could be proven end to end before anything harder.
- One intent: NTP server and syslog collector are both
192.168.100.10. - Everything runs on a single Ubuntu control node inside the lab. Ansible and its dependencies are locked with uv, so the toolchain version is in git alongside the model.
- Versions: ansible-core 2.21.4, ansible.netcommon 8.6.2, arista.eos 12.2.0, cisco.nxos 11.2.0, paramiko 5.0.0.
Nothing here says the same code is safe on a production network. It says what a lab found before a production network had the chance to.
The layout: two repositories#
~/iac/model/ # what should be true — inventory, intent, templates, playbooks
~/iac/backups/ # what the devices actually say — one file per deviceThey are separate on purpose. Backups contain password hashes and must never land in the model repo. Keeping them apart also makes git log on the backups repo an audit trail of each device’s config, as it was at each collection:
5abaa54 NXOS, Cumulus01: remediated, converged
57fb18d Arista01: syslog remediated, converged
0201415 brownfield: hand-made legacy NTP/syslog state
f750f8d baseline: eos, nxos, cumulus“At each collection” is the limit. A change that is made and reverted between two collections never reaches git. That happened to me during the first run of the drift test below, and it’s why that test was repeated.
Setting up the Ansible project#
Toolchain#
Ansible is a Python package, so its version is a dependency like any other. I manage it with uv. uv.lock records exact versions and goes into git. .venv/ stays out of git, and uv sync rebuilds it identically on another machine:
uv init --bare
uv add ansible paramiko[project]
name = "model"
version = "0.1.0"
requires-python = ">=3.12"
dependencies = [
"ansible>=14.4.0",
"paramiko>=5.0.0",
]Every command then runs as uv run ansible-playbook ..., or plain ansible-playbook after activating .venv. If ansible comes back “not found”, the shell isn’t in the venv. The fix is uv run, not apt install ansible-core, which would install a second Ansible that uv.lock doesn’t control.
ansible.cfg#
[defaults]
inventory = inventory/lab.yml
host_key_checking = False
vault_password_file = ~/.vault_pass
stdout_callback = default
callback_result_format = yaml
log_path = ./logs/ansible.log
[persistent_connection]
command_timeout = 60host_key_checking = Falseis for the lab only, where devices are rebuilt and their keys change.vault_password_filepoints outside the repo, so the key that decrypts the vaults never goes into git.callback_result_format = yamlreplaces the oldstdout_callback = yaml. That plugin has been removed fromcommunity.general, and a config that still names it stops every run with an error. My first replacement wasresult_format = yaml. That’s the callback option’s name, but not its INI key, so Ansible ignored it without a word and the output stayed JSON. I only noticed because a reviewer pointed out the key; the JSON output had been in front of me the whole time.- Without
log_path, Ansible writes to the terminal and nowhere else. I learned that when I wanted the results of an earlier run and they didn’t exist.
Inventory and variables#
all:
children:
eos:
hosts:
Arista01: {ansible_host: 192.168.100.11}
nxos:
hosts:
NXOS: {ansible_host: 192.168.100.12}
cumulus:
hosts:
Cumulus01: {ansible_host: 192.168.100.13}Groups are platforms. Every variable lives in the group_vars folder at the level where it’s true:
inventory/group_vars/
all/vars.yml # same on every device: username, password pointer
all/intent.yml # what must be true (the model)
eos/vars.yml # how to reach an EOS device, and its syntax facts
eos/vault.yml # the Arista password, encrypted
nxos/... # same pair per platform
cumulus/...ansible_user: darren
ansible_password: "{{ vault_ansible_password }}"ansible_network_os: arista.eos.eos
ansible_connection: ansible.netcommon.network_cli
ansible_network_cli_ssh_type: paramiko
platform: eos
mgmt_vrf: default
owned_regex: '^(ntp server |logging (vrf \S+ )?host )'
negate_regex: '^'
negate_with: 'no '
remediation_footer: []
ansible_become: true
ansible_become_method: enableThe EOS user logs in at Arista01>, so become + enable is what gets show running-config to work. (It needs privileged mode.)
Each vault.yml holds exactly one line, vault_ansible_password: '<that platform's password>', and is encrypted with ansible-vault create. The vault_ prefix is a convention. The readable pointer in all/vars.yml shows which variable is secret and where it comes from, and only the value is encrypted. Two checks confirm it resolves, and neither one prints a secret:
# key names only, values masked
for f in inventory/group_vars/*/vault.yml; do
echo "== $f"; ansible-vault view "$f" | sed -E 's/(: *).*/\1***/'
done
# what each host will actually log in with, by length
ansible all -m ansible.builtin.debug \
-a "msg='user={{ ansible_user }} pwlen={{ ansible_password | length }}'"What Ansible doesn’t check#
Ansible accepts any variable name or config key, so a misspelled one isn’t an error. It’s just a setting nothing reads. In this project that happened four times:
ansibel_hostin the inventory. Caught when I reviewed the file, before it ran.nsibel_password: "{{ value_ansible_password }}"inall/vars.yml, with two typos in one line.ansible_become_mothodin the EOS vars.result_formatinstead ofcallback_result_formatinansible.cfg. The output silently stayed JSON.
The last three sat in a working system. The vaults still held ansible_password directly, and a platform’s group_vars overrides all, so every login worked and every device converged. I only found them while writing this article. A fifth typo was in a value, not a name (mgmt_vrf: defatult, point 1 below). No tool could reject that one, because none of them knows which VRF names are valid.
Fixing them caused a regression that’s worth more than the typos. I assumed the EOS user landed in privileged mode, so I deleted both become lines. Collection then failed with privileged mode required: ansible_become: true had been doing the work all along. My checks after that edit had passed, because show hostname doesn’t need privileged mode and the offline diff never connects. Neither one exercised what the edit changed.
That class of bug is what the CI in part 2 is designed to catch.
Collect: read-only#
The backup playbook only ever sends show commands. It uses ansible.netcommon.cli_command for EOS and NX-OS, and nv config show -o commands for Cumulus, which returns NVUE’s applied config as a list of nv set lines. It does not use a *_config module with backup: yes.
That makes it read-only because every command it sends only retrieves configuration, not because the module can’t write. cli_command will send whatever it’s given, configure included. The platform *_command modules don’t change that: the cisco.nxos.nxos_command documentation itself shows configure terminal followed by a config change. The read-only boundary is the list of commands in the playbook, so that list is what has to be reviewed.
- name: Collect running config (read-only)
hosts: eos,nxos,cumulus
gather_facts: false
vars:
backup_dir: "{{ lookup('env', 'HOME') }}/iac/backups"
tasks:
- name: EOS / NX-OS - show running-config
ansible.netcommon.cli_command:
command: show running-config
register: netos
when: "'cumulus' not in group_names"
- name: Cumulus - NVUE applied config as commands
ansible.builtin.command: nv config show -o commands
register: nvue
changed_when: false
when: "'cumulus' in group_names"
- name: Write backup file
ansible.builtin.copy:
content: "{{ (nvue.stdout if 'cumulus' in group_names else netos.stdout) | regex_replace('(?m)^!(Time|Running configuration last done at):.*\\n', '') }}\n"
dest: "{{ backup_dir }}/{{ inventory_hostname }}"
mode: "0600"
delegate_to: localhostThe two collection tasks register different variable names on purpose. A skipped task still overwrites the variable it registers, so if both used one name, the skipped task would replace the real output with a “skipped” result.
The test for collection is not “files exist”. It is run it twice and git status stays clean. Any line that changes on every collection, such as a timestamp header, turns every collection into a commit and every comparison with an earlier backup into a diff. NX-OS starts its running-config with exactly that kind of header:
!Command: show running-config
!Running configuration last done at: Wed Sep 30 05:41:21 2026
!Time: Wed Sep 30 05:50:58 2026!Time: is the moment of the show, so it’s different on every collection. last done at changes whenever anything is configured, including changes the model doesn’t own. The regex_replace strips both. I added it before the first collection, so I didn’t see the noise happen. The output above, from a later run, shows it would have.
This is about the backup files and their git history, not the drift report. The drift report only looks at owned lines, and no owned regex matches these headers, so they could never have caused false NTP or syslog drift. What stripping them buys is a stable history: a new commit means the device config changed, and git diff against an earlier backup shows only real changes. The drift test below relies on that.
The same \\n escaping that later failed in a join (see Smaller ones) works here, because a regex treats a literal backslash-n as a newline.
Model: what, how, and syntax are three different things#
The intent is one file and says nothing vendor-specific:
# group_vars/all/intent.yml
ntp_servers:
- 192.168.100.10
syslog_hosts:
- 192.168.100.10How to reach those servers — which VRF — is a per-platform fact, kept next to each platform’s connection settings:
# group_vars/eos/vars.yml mgmt_vrf: default
# group_vars/nxos/vars.yml mgmt_vrf: management
# group_vars/cumulus/vars.yml mgmt_vrf: mgmtA template per platform turns the two into syntax. The same intent, one NTP server reached through the platform’s management VRF, puts the VRF in a different place on each platform. On the EOS device in this lab the management VRF is the default VRF, so it doesn’t appear at all:
| Platform | Rendered line |
|---|---|
| EOS (non-default VRF, not tested here) | ntp server vrf MGMT 192.168.100.10 |
| EOS (default VRF, this lab) | ntp server 192.168.100.10 |
| NX-OS | ntp server 192.168.100.10 use-vrf management |
| Cumulus NVUE | nv set service ntp mgmt server 192.168.100.10 |
Before the VRF, after the address, or part of the command path. None of these is harder than the others; they are just all different, and the model has to know. The three templates:
{% for s in ntp_servers %}
ntp server {% if mgmt_vrf != 'default' %}vrf {{ mgmt_vrf }} {% endif %}{{ s }}
{% endfor %}
{% for s in syslog_hosts %}
logging {% if mgmt_vrf != 'default' %}vrf {{ mgmt_vrf }} {% endif %}host {{ s }}
{% endfor %}{% for s in ntp_servers %}
ntp server {{ s }} use-vrf {{ mgmt_vrf }}
{% endfor %}
{% for s in syslog_hosts %}
logging server {{ s }} use-vrf {{ mgmt_vrf }}
{% endfor %}{% for s in ntp_servers %}
nv set service ntp {{ mgmt_vrf }} server {{ s }}
{% endfor %}
{% for s in syslog_hosts %}
nv set service syslog {{ mgmt_vrf }} server {{ s }}
{% endfor %}The NX-OS template is the corrected version. Its first version had a 5 in the logging line, and point 2 below explains why it came out.
A render playbook writes one file per device to rendered/. Every task in it is delegated to localhost, so rendering never needs a device. I tested that by pointing every ansible_host at an address the control node can’t reach (-e ansible_host=203.0.113.1, from the TEST-NET-3 documentation range). All three renders still came back ok.
Diff offline: owned lines only#
Each platform declares which lines the model owns with one regex:
owned_regex: '^(ntp server |logging (vrf \S+ )?host )' # EOS
owned_regex: '^(ntp server |logging server )' # NX-OS
owned_regex: '^nv set service (ntp|syslog) \S+ server \S+$' # CumulusThe Cumulus regex is anchored to server entries only. My first version, '^nv set service (ntp|syslog) ', matched every setting under those two services, while the template renders only servers. Any other NTP or syslog setting on the device would have counted as extra, and the diff would have generated an nv unset for it. It never happened in this lab, because the only settings under those services were servers, but a reviewer pointed out that the claimed scope and the regex didn’t agree.
The diff takes the owned lines from the backup and the rendered lines from the template, and reports two sets per device: missing (rendered, not on the device) and extra (owned, on the device, not rendered). Everything the regex does not match is ignored: counted as not mine, and never touched.
Two rules are enforced in code rather than in a runbook:
- An empty intent refuses to run. If
ntp_serverswere emptied by mistake, every NTP line on every device would become “extra” and the generated remediation would delete them all. Anassertat the top of the playbook stops that. - Adds before removals. On EOS and NX-OS, each line takes effect as it’s sent. For replacements involving distinct server entries, the generated commands add the replacement before removing the old entry. That ordering doesn’t help when the two lines are different spellings of the same server: point 2 below is a case where the “add” was a no-op and the “remove” would have deleted the only syslog server. On Cumulus the order doesn’t matter, because NVUE stages every
nv set/nv unsetand applies them together atnv config apply. The rule is still applied there so the playbook has one behavior.
Remediation is generated, never hand-written. Missing lines are sent as-is; extra lines are negated per platform: no prepended on EOS and NX-OS, nv set rewritten to nv unset on Cumulus, followed by nv config apply -y. Two variables per platform, negate_regex and negate_with, drive this, so the playbook itself has no vendor logic in it:
- name: Render first
ansible.builtin.import_playbook: render.yml
- name: Offline drift report (never touches a device)
hosts: eos,nxos,cumulus
gather_facts: false
vars:
model_dir: "{{ playbook_dir }}/.."
backup_dir: "{{ lookup('env', 'HOME') }}/iac/backups"
tasks:
- name: Refuse an empty intent
ansible.builtin.assert:
that:
- ntp_servers | length > 0
- syslog_hosts | length > 0
fail_msg: "An empty intent list would delete every server in that domain. Refusing."
- name: Load owned lines from backup, and intended lines from render
ansible.builtin.set_fact:
actual: "{{ lookup('file', backup_dir ~ '/' ~ inventory_hostname).splitlines() | select('match', owned_regex) | list }}"
wanted: "{{ lookup('file', model_dir ~ '/rendered/' ~ inventory_hostname).splitlines() | reject('equalto', '') | list }}"
- name: Compute drift
ansible.builtin.set_fact:
missing: "{{ wanted | difference(actual) | sort }}"
extra: "{{ actual | difference(wanted) | sort }}"
- name: Build remediation (adds first, then removals)
ansible.builtin.set_fact:
remediation: "{{ missing + (extra | map('regex_replace', negate_regex, negate_with) | list) }}"
- name: Write remediation file
ansible.builtin.template:
src: "{{ model_dir }}/templates/remediation.j2"
dest: "{{ model_dir }}/remediation/{{ inventory_hostname }}.cfg"
mode: "0644"
delegate_to: localhost
- name: Drift summary
ansible.builtin.debug:
msg: "owned={{ actual | length }} missing={{ missing | length }} extra={{ extra | length }}"{% for line in remediation %}
{{ line }}
{% endfor %}
{% if remediation %}
{% for line in remediation_footer %}
{{ line }}
{% endfor %}
{% endif %}The footer holds the commit step on platforms that need one, such as nv config apply -y on Cumulus. It’s only written when there’s something to commit, so a converged device gets an empty file.
The brownfield state being adopted looked like this. Once the typo in point 1 below was fixed, the report matched it:
| Device | Before | Report |
|---|---|---|
| Arista01 | NTP .10, no syslog | syslog missing |
| NXOS | NTP .10, no syslog | syslog missing |
| Cumulus01 | four vendor pool NTP servers, syslog .10 | NTP .10 missing, four pool servers extra |
The Cumulus row is the realistic one. Nobody chose those pool servers; they shipped with the image and were never adopted.
Apply: one device, read first, prove it after#
The apply playbook runs the change procedure end to end, and nothing in it is optional. It takes a fresh backup in the same run and diffs against it. It shows the exact lines it will send and waits for a person to confirm. It pushes to one device at a time. Then it collects and diffs again:
- name: Fresh backup (same run)
ansible.builtin.import_playbook: backup.yml
- name: Diff against that backup
ansible.builtin.import_playbook: diff.yml
- name: Apply remediation (lab only, one device at a time)
hosts: eos,nxos,cumulus
gather_facts: false
serial: 1
vars:
model_dir: "{{ playbook_dir }}/.."
tasks:
- name: Load remediation
ansible.builtin.set_fact:
fix_lines: "{{ lookup('file', model_dir ~ '/remediation/' ~ inventory_hostname ~ '.cfg').splitlines() | reject('equalto', '') | list }}"
- name: No drift, nothing to send
ansible.builtin.meta: end_host
when: fix_lines | length == 0
- name: Exactly what will be sent
ansible.builtin.debug:
var: fix_lines
- name: Confirm
ansible.builtin.pause:
prompt: "Send the lines above to {{ inventory_hostname }}? Enter = yes, Ctrl+C then A = abort"
- name: Push (EOS)
arista.eos.eos_config:
lines: "{{ fix_lines }}"
save_when: modified
when: "'eos' in group_names"
- name: Push (NX-OS)
cisco.nxos.nxos_config:
lines: "{{ fix_lines }}"
save_when: modified
when: "'nxos' in group_names"
- name: Push (Cumulus)
ansible.builtin.command: "{{ item }}"
loop: "{{ fix_lines }}"
when: "'cumulus' in group_names"
- name: Re-collect after the change
ansible.builtin.import_playbook: backup.yml
- name: Verify convergence (fails the run on any drift)
ansible.builtin.import_playbook: verify.ymlThe last step used to be a plain re-diff. It printed missing and extra, but the run finished green even when they weren’t zero. So convergence was reported, not enforced. Now it’s a separate playbook that fails:
- name: Diff against the current backups
ansible.builtin.import_playbook: diff.yml
- name: Convergence gate - fail if anything is missing or extra
hosts: eos,nxos,cumulus
gather_facts: false
tasks:
- name: Assert zero drift
ansible.builtin.assert:
that:
- missing | length == 0
- extra | length == 0
fail_msg: "Not converged: missing={{ missing }} extra={{ extra }}"A gate that has only ever passed hasn’t been tested. So I added a second NTP server to the intent without touching any device, and ran verify.yml:
fatal: [Arista01]: FAILED! =>
msg: 'Not converged: missing=[''ntp server 192.168.100.11''] extra=[]'
fatal: [NXOS]: FAILED! =>
msg: 'Not converged: missing=[''ntp server 192.168.100.11 use-vrf management''] extra=[]'
fatal: [Cumulus01]: FAILED! =>
msg: 'Not converged: missing=[''nv set service ntp mgmt server 192.168.100.11''] extra=[]'
exit code: 2It printed one missing line per platform, each in that platform’s syntax, and a non-zero exit code. After I restored the intent, the same run exited 0. verify.yml only reads the backup files in the working directory and the rendered intent, so it needs no access to the lab. That’s also its limit. It checks a snapshot: a pass means the model matches those files. It doesn’t check whether they’re committed, or how old they are. Inside apply.yml, the collection just before it provides fresh evidence, so a pass there does describe the devices. Run on its own against old backups, as a CI job would, it can only show that a proposed change matches the last collected state, not what the devices look like now. Part 2 has to design around that difference.
What failed, and where it showed up#
1. A typo that generated a command to delete the only NTP server#
My EOS vars said mgmt_vrf: defatult. The template checks for default, so it rendered ntp server vrf defatult 192.168.100.10. The diff then did exactly what it should: the device’s real line, ntp server 192.168.100.10, no longer matched the model, so it was extra. The generated remediation was:
ntp server vrf defatult 192.168.100.10
logging vrf defatult host 192.168.100.10
no ntp server 192.168.100.10Add two lines in a VRF that does not exist, then remove the one NTP server that works. Nothing in the tooling was wrong; the model was.
This one was caught before any write path existed. The apply playbook hadn’t been written yet. The remediation files were only generated, and reading them before going further was part of the step. The Confirm pause in apply.yml exists because of this. It can’t make anyone read, but it stops the run and puts the exact lines in front of whoever is about to press Enter.
2. NX-OS does not store what you send it#
The NX-OS template rendered logging server 192.168.100.10 5 use-vrf management. The device accepted it. The re-collected running-config said:
logging server 192.168.100.10 use-vrf managementThe severity field was gone — I assume because 5 is the default, though I have not tested a non-default value to confirm it. The diff now saw the rendered line as missing and the device’s line as extra, and nothing in the model would ever change that. The generated remediation would have re-added a line that was already there and then removed the syslog server. I read it and didn’t push it.
This is the one failure in this post that only step 5 could have shown. The push succeeded and no error appeared anywhere; the model and the device simply disagreed about what “the same line” looks like. The fix was in the template, not the device: render the form the device stores, not the form you would type. Defaults it drops, abbreviations it expands and keyword order it normalizes all break convergence the same way, and the only way I know to find them is to push to a lab device and re-collect.
3. Undo on a tree is not the reverse of do#
For the drift test I added a static route on Cumulus, a change outside the model, which the diff should ignore. It did. Removing it by hand afterwards was the surprise.
One nv set vrf default router static 10.99.99.0/24 via ... had created the parent nodes it needed. Unsetting the route left an empty parent behind, and unsetting that left the next one:
nv set vrf default router static # after removing the route
nv set vrf default router # after unsetting 'static'
nv set vrf default # after unsetting 'router'
# after unsetting 'vrf default': cleanIt took four unsets, checking the config after each, before the backup matched the baseline again. EOS and NX-OS did not behave this way in this lab; no <line> removed the line. NVUE’s config is a tree, and on Cumulus VX 5.9.1 at least, removing a leaf did not prune the parents it created.
It matters for the model too, though not in the way I first expected. Suppose an intent removes the last server in a VRF, and nv unset service ntp mgmt server X leaves an empty nv set service ntp mgmt behind. The owned regex now only matches server <value> lines, so the drift report would ignore that container. The diff would converge while leftover config stayed on the device. Only a comparison of the whole backup file with an earlier one would show it. I have not tested that case yet. What this does justify is testing restore behavior before trusting line-by-line undo on a tree-structured platform. One experiment is to restore a whole known-good config (nv config replace, or an earlier NVUE revision) and compare the result with the baseline. Whether that’s the right approach here is still open, because the model owns only part of the config and a whole-config restore would overwrite everything else.
4. The same drift, a different file every run#
The first Cumulus remediation listed the four unset lines in a different order each run: 1, 2, 3, 0, then 3, 0, 1, 2. Ansible’s difference filter doesn’t preserve order. When the items can be hashed, it works on sets, and Python randomizes string hashing per process, so each run can come out in a different order. For NTP servers that is harmless, but a generated file that changes when nothing has changed is noise in every comparison after it. Sorting both sets fixed it. That is only correct because this domain is order-independent; an ACL or prefix-list could not be sorted, and that is the reason the next domain will be one.
Smaller ones#
- Old SSH, new library. Before moving to these images, the lab ran Cisco vIOS 15.9, which offered only SHA1 key exchange. paramiko 5.0.0 offers none. Ansible failed with
no acceptable kex algorithm. I pinned paramiko below 4 and then retired the vIOS nodes before confirming the pin worked. The lesson stands anyway: an old device can hold the control node’s tooling back. - A literal
\n. Joining the remediation lines withjoin('\n')inside an inlinecopycontent:wrote a literal backslash-n on the ansible-core version listed above, in both double-quoted and block-scalar YAML. I did not dig into why. Writing the file through thetemplatemodule, the same way the rendered configs are written, removed the problem.
The drift test#
With all three devices converged, I made three changes by hand to see whether the detector found what someone else might do:
| Change | Drift summary | Generated remediation |
|---|---|---|
Arista01: add ntp server 192.168.100.99 | owned=3 missing=0 extra=1 | no ntp server 192.168.100.99 |
| NXOS: remove the syslog server | owned=1 missing=1 extra=0 | logging server 192.168.100.10 use-vrf management |
| Cumulus01: add a static route (outside the model) | owned=2 missing=0 extra=0 | (empty) |
The Cumulus row is the one I cared about most. The route was in the backup, and the backups repo recorded it: the “injected” commit changed all three files. But the model doesn’t own it, so the diff didn’t report it and nothing was generated to remove it. A change outside the model stays visible in the history, and the tool never touches it.
apply.yml then showed each device’s single line at the pause, pushed it, and re-collected. Both devices came back at owned=2 missing=0 extra=0. I removed the Cumulus route by hand, collected again, and compared the backups with the converged commit from before the test:
git -C ~/iac/backups diff 5abaa54 --statIt printed nothing: the three normalized backup files (collected, with the NX-OS timestamp headers stripped) matched the pre-test commit exactly. The test sits in the backups history as two commits, one with the drift injected and one with it remediated.
There was one warning worth noting. On both the EOS and NX-OS pushes, the modules warned that the input lines should match how they appear in the running config, or idempotency and the diff can’t be guaranteed. It’s point 2 again, this time from the vendor modules themselves, which appear to compare lines as text too.
Why not resource modules?#
Ansible already has a structured answer to “compare the model with the device”: resource modules such as cisco.nxos.nxos_logging_global and arista.eos.eos_ntp_global. With state: parsed they turn config text into structured data, and with state: overridden they converge a device to it. So the obvious question is whether they’d have made point 2 go away.
I ran their parsers against the same backups. parsed works fully offline: every run below targeted the fake 203.0.113.1 and succeeded. It only requires the connection type to be network_cli.
The direct test is to feed the NX-OS parser the two forms of the syslog line, the one my first template wrote and the one the device stored:
| Input line | nxos_logging_global parsed |
|---|---|
logging server 192.168.100.10 5 use-vrf management | host: 192.168.100.10, severity: notification, use_vrf: management |
logging server 192.168.100.10 use-vrf management | host: 192.168.100.10, use_vrf: management |
Severity 5 becomes severity: notification. When severity is left out, the field is absent, not filled in as the default. The two forms are still different after parsing. The EOS parser does the opposite. My backup line is logging host 192.168.100.10 with no port, and eos_logging_global returned port: 514, filling in a default the device doesn’t show.
So in this lab, parsing into structured data didn’t make defaults disappear from the parsed output. One parser keeps a default the device drops, and the other adds a default the device hides. That’s a statement about parsed only. It doesn’t show how the modules behave when converging: state: overridden runs its own comparison, which may normalize defaults, and I haven’t tested it. Until I do, the most I can say is that the parsed data shows you still need to know each platform’s defaults.
The other reason for the text-based pipeline is uniformity. EOS and NX-OS have resource modules for both domains. Cumulus NVUE is a tree reached through nv commands or a REST API. It’s a different model, and I haven’t evaluated the nvidia.nvue collection. One owned-line diff over collected text works the same for all three, and its output is plain text I can read and put in git. The cost is that the text itself is the contract, which is exactly what points 1 and 2 are about.
What this does not show#
- One domain, and the easy one. NTP and syslog lines are flat and order-free. Nothing here covers ordered policy, nested interface config or anything whose removal can cut you off.
- Anything the model doesn’t own is invisible to the drift report. Checking the NX-OS output while writing this, I found a second admin account,
darrem, next todarren. It hadrole network-adminand a matching SNMPv3 user. It was a typo from setting up the lab, and I’ve since removed it. It had been in the backups since the baseline commit, and every drift report saidmissing=0 extra=0, because the model owns NTP and syslog lines and nothing else. The backups made it findable, but only because someone happened to look. Owning a domain is what turns “findable” into “reported”. - One device per platform, virtual images. Behaviors here are observed on these image versions, not general statements about the platforms.
- No restore test on Cumulus. Given point 3, that is the next one to prove.
- Manually triggered, and local. Drift is only detected when I run it, and both repositories live on one control node. Moving them to a remote Git repository with review and CI is the next part of this series.
Conclusion#
For this lab, configuration convergence required a fresh collection with no missing or extra owned lines. “The playbook ran green” and “the device accepted the commands” weren’t enough: the NX-OS failure passed both, and only showed up when the device was asked afterwards what it had stored.
That’s a claim about configuration, not about service. A converged backup shows that the owned lines match the model. It doesn’t show that NTP is synchronized or that syslog messages arrive at the collector. Checking those is a separate test, and this post doesn’t cover it.
Next in this series: Part 2 puts the model and the backups on Git properly: a remote repository, changes reviewed before they’re merged, and CI that runs the syntax check, lint and offline diff on every change.