Driving network devices from a playbook is popular because it removes the two things people dislike about scripting: there is no connection handling to write, and the description of what you want is separate from the mechanics of getting there. Both of those are real and both make the tool worth using.
What they also do is hide the mechanics well enough that a playbook can express something considerably more destructive than its author intended. One word in a task decides whether the tool adds your configuration to what is there or makes the device match your data and removes everything else. Both spellings look equally routine, and the second one run against a partial data set is a configuration eraser.
This article covers how the tool reaches a device at all, the difference between the module that pushes commands and the modules that describe state, what the word "changed" actually means in the output, how to bound what a run can affect, and the playbooks that cause damage. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.
How Does It Reach a Device?
What has to be right?
An inventory entry naming the device, a connection type suited to network equipment rather than to servers, the platform's name so the right module implementations are used, credentials, and a statement that privileged mode is required. Getting the connection type or the platform name wrong produces a failure that reads like an authentication problem, because the tool tries to behave as though the device were a server and the device does not cooperate.
A Deeper Dive into Connectivity
Why the connection type matters
The default behaviour assumes a general-purpose operating system with a shell, a file system it can copy a module into, and an interpreter to run it. A network device has none of those in the expected form, so the tool has a separate connection type that keeps the logic on the control machine and sends commands over the session.
Omitting it produces a failure about being unable to execute something on the device, which reads as a permissions problem and is a connection type problem.
# Inventory: the four settings that make it work at all
all:
children:
ios:
hosts:
R1: {ansible_host: 10.1.1.1}
R2: {ansible_host: 10.1.1.2}
vars:
ansible_connection: ansible.netcommon.network_cli
ansible_network_os: cisco.ios.ios
ansible_user: "{{ lookup('env', 'NETUSER') }}"
ansible_password: "{{ lookup('env', 'NETPASS') }}"
ansible_become: true
ansible_become_method: enable
ansible_command_timeout: 60
Privileged mode
Configuration requires it and the session does not start there unless told. A playbook without the escalation settings connects successfully, runs a read-only task fine, and fails on anything that configures — which makes it look like the configuration task is wrong rather than the connection.
Where the account already lands in privileged mode, the setting is harmless. Including it unconditionally is simpler than deciding per device.
Fact gathering
The default behaviour of collecting information about the host assumes a server and either fails or wastes time on a network device. Turning it off in the play and using the platform's own information-gathering module where facts are needed is the correct arrangement.
This is one line and its absence is a common cause of a playbook that takes far longer than it should, or that fails at the very first step for reasons unrelated to the task.
# The play header for network devices
- name: Configure interfaces
hosts: ios
gather_facts: false
connection: ansible.netcommon.network_cli
tasks:
- name: Collect device facts properly
cisco.ios.ios_facts:
gather_subset: [hardware, interfaces]
- name: Show what we learned
ansible.builtin.debug:
msg: "{{ ansible_net_version }} on {{ ansible_net_model }}"
Credentials
Not in the inventory file as plain text. Either from the environment, or from the tool's own encrypted storage, or from a secret manager. The inventory is committed to version control and a password in it is a password in version control.
The encrypted storage is the lowest-friction option that is actually acceptable, and it integrates with the inventory directly.
Timeouts
A device that is slow to answer — a large configuration, a busy processor — will exceed the default command timeout and produce a failure that looks like a connectivity problem. Raising it for playbooks that read large outputs is routine.
The persistent connection also has its own idle timeout, and a play with a long gap between tasks against the same host can lose its session. Both are settings rather than problems once you know they exist.
| Setting | Omitting it produces |
|---|---|
| Network connection type | Failures that look like authentication |
| Platform name | Wrong module implementation, odd errors |
| Privilege escalation | Reads work, configuration fails |
| Fact gathering disabled | Slow starts or an immediate failure |
| Command timeout raised | Large outputs fail as connectivity problems |
Command Pusher or Resource Module?
What is the difference?
The command pusher takes lines of configuration and sends them, deciding whether to bother by matching the text against the running configuration. That match is approximate, so it sometimes reports a change that did not happen and sometimes misses one that should. A resource module takes a description of the desired state as structured data, reads what the device currently has, computes the difference and sends only that. The second is genuinely idempotent; the first is idempotent-ish.
A Deeper Dive into the Two Kinds
What the command pusher actually does
It checks whether the lines you supplied appear in the running configuration within the context you named, and if they do not, it enters that context and sends them. The check is textual, so a line that is present in a slightly different form — a different abbreviation, a value the device normalised — is not recognised.
The result is a task that reports a change on every run without changing anything, which is noise, or a task that reports no change when the configuration has drifted, which is worse.
# The command pusher: you supply lines, it matches text
- name: Configure an interface description
cisco.ios.ios_config:
parents: "interface GigabitEthernet0/1"
lines:
- "description Uplink to core"
- "ip address 10.1.12.1 255.255.255.252"
match: line
save_when: modified
!
# Works anywhere, for anything. Idempotency is textual and approximate.
What a resource module does
It reads the relevant part of the configuration, parses it into structured data, compares that against the data you supplied, computes the commands needed to reconcile the two, and sends them. If nothing is needed, it sends nothing and reports no change.
That is a real comparison rather than a text match, so the report is trustworthy and re-running is genuinely free. Its limitation is coverage: it handles what the module models and nothing else.
# A resource module: you describe the state, it works out the commands
- name: Ensure interface configuration matches
cisco.ios.ios_l3_interfaces:
config:
- name: GigabitEthernet0/1
ipv4:
- address: 10.1.12.1/30
- name: Loopback0
ipv4:
- address: 10.255.0.1/32
state: merged
When to use which
A resource module wherever one covers what you are configuring, because the idempotency is real and the data is easier to review than a list of commands. The command pusher for anything not covered, which on a large platform is a lot of things.
Mixing both in one playbook is normal and fine. What is worth avoiding is using the command pusher for something a resource module covers, because that gives up the guarantee for no benefit.
The state values, and the one to be careful with
Adding your data to what is there. Rewriting the objects you named while leaving others alone. Making the device match your data entirely, which removes every object of that kind that your data does not mention. Removing what you named. And three read-only values that change nothing.
The third one is the dangerous one and it is spelled almost the same as the second. Run against a data set listing three interfaces on a switch with forty-eight, it removes the configuration from forty-five of them.
The read-only states
One reads the device and returns the structured data, which is how you discover what the module's data model looks like for something already configured. One takes your data and produces the commands it would send, without touching a device at all. One takes a configuration text and parses it into the module's data model.
The second is the most useful and the least used. It converts a playbook into a review artefact in one run with no device involved, which is the cheapest possible safety step.
# Produce the CLI without touching anything
- name: Show what would be sent
cisco.ios.ios_l3_interfaces:
config: "{{ interface_config }}"
state: rendered
register: preview
- ansible.builtin.debug:
var: preview.rendered
!
# And learn the data model from a configured device
- name: Read the current state as data
cisco.ios.ios_l3_interfaces:
state: gathered
register: current
Templates and the command pusher
A template rendering a block of configuration, fed to the command pusher, is a common and workable pattern for things no resource module covers. The template's output is text and inherits every property of generated text, including that nothing validates it.
Where that pattern is used, rendering to a file and reviewing it before the run is the safety step, exactly as it is for any generated configuration.
| Need | Use | Why |
|---|---|---|
| Something a module covers | Resource module | Real idempotency, reviewable data |
| Anything else | Command pusher | Covers everything, approximately |
| Read state as data | gathered | Learn the model from a real device |
| See the commands first | rendered | No device involved at all |
| Read-only checks | The command runner | Never reports a change |
What Does "Changed" Actually Mean?
Can it be trusted?
It depends entirely on the module. From a resource module it means a real difference was computed and applied, and it is trustworthy. From the command pusher it means a text match failed, which may or may not correspond to an actual difference. From the read-only command runner it never appears at all, regardless of what the commands returned. A summary line counting changes therefore means different things for different tasks in the same play.
A Deeper Dive into the Report
The perpetually changing task
A command pusher task whose lines never match the device's rendering of them reports a change on every run and sends the same commands every time. The configuration is correct throughout; the report is wrong.
Common causes: a value the device normalises to a different form, a command that appears in the configuration under a different parent, or an abbreviation. The fix is to write the lines exactly as the device renders them, which means reading them from the device rather than from documentation.
The task that misses a difference
The mirror image, and the more serious one. A device whose configuration has drifted in a way the text match does not notice reports no change, so the playbook run confirms compliance that does not exist.
This is the argument for resource modules stated most sharply: a compliance run built on approximate matching reports compliance approximately, which for a compliance run is the one property it must not have.
Reading the difference rather than the summary
The diff output shows what was actually sent, per device, which is the only trustworthy view. A summary line saying three hosts changed says nothing about what changed on them.
Running with the difference displayed, always, converts the output from a count into a record. It is one flag and it makes the run reviewable.
# Always. The summary line is not enough.
$ ansible-playbook interfaces.yml --check --diff --limit R1
!
$ ansible-playbook interfaces.yml --diff --limit R1
!
# The diff shows the commands per host, which is the record.
Check mode
The tool can run a play and report what it would do without doing it. For resource modules that is accurate, because the difference is computed the same way either way. For the command pusher it is approximate for the same reason its change reporting is.
It is still worth using on both, because approximate is a great deal better than nothing when the alternative is applying a change blind.
Saving the configuration
A change applied to the running configuration is not saved unless something saves it. The command pusher has a setting for this and it is not on by default; resource modules generally do not save either.
A playbook that changes configuration and never saves produces devices that revert on reload, which is discovered weeks later. The save belongs either as a setting on the tasks that change things or as a final task in the play.
# Nothing saves automatically. Two ways to do it.
- name: Configure, and save only if something changed
cisco.ios.ios_config:
lines: ["description managed by automation"]
parents: "interface Loopback0"
save_when: modified
!
# Or once, at the end of the play
- name: Save if anything changed
cisco.ios.ios_config:
save_when: always
when: ansible_play_hosts_all | length > 0
Handlers
A task that should run only if something else changed — saving, or restarting a process — is expressed as a handler triggered by the change notification. That is the idiomatic arrangement and it depends on the change reporting being accurate.
With a task that reports a change every run, the handler fires every run. That turns an inaccurate report into an action, which is how a noisy playbook becomes a playbook that writes to storage on every scheduled execution. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.
| Module | changed means | Trust the report |
|---|---|---|
| Command runner | Never set | n/a — read only |
| Command pusher | A text match failed | Approximately |
| Resource module | A computed difference was applied | Yes |
| In check mode | Would have done the above | Same caveats |
How Do You Bound What a Run Affects?
What are the controls?
Four, and they compose. Preview mode with the difference shown, which changes nothing. A limit restricting the run to named hosts, so the first real run touches one device. Batching, so a run across many devices proceeds a few at a time and stops if a batch fails. And a failure policy deciding whether one host's failure stops the play. Used together they turn a playbook run from an event into a procedure.
A Deeper Dive into Blast Radius
The limit
Restricting a run to named hosts or a pattern. The first application of any new or modified playbook should use it with a single host, and the second with a handful.
It is also the protection against an inventory that grew. A playbook written when the group held ten devices runs against the two hundred it holds now, and the limit is what makes that a deliberate decision rather than a discovery.
# The sequence, four commands
$ ansible-playbook site.yml --check --diff --limit R1 # see it
$ ansible-playbook site.yml --diff --limit R1 # do one
$ ansible-playbook site.yml --diff --limit 'branch[0:4]' # do a few
$ ansible-playbook site.yml --diff # the rest
Batching
By default the tool works on many hosts in parallel and completes each task across all of them before moving to the next. Batching changes that to completing the whole play on a subset before starting the next subset.
For a configuration change that matters, batching is what limits a bad change to one batch. Combined with a failure policy that stops the play when a batch fails badly enough, it means a mistake reaches a fraction of the estate rather than all of it.
# A few at a time, and stop if a batch goes wrong
- name: Roll out the change
hosts: branch_routers
gather_facts: false
serial: 5
max_fail_percentage: 20
any_errors_fatal: false
tasks:
- name: Apply
cisco.ios.ios_config:
src: templates/branch.j2
save_when: modified
The failure policy
By default a host that fails is removed from the play and the others continue. For an independent change that is right; for a change where a partial application across the estate is worse than none, stopping everything on the first failure is right.
Deciding which applies is a property of the change rather than a default to accept. A change that must be everywhere or nowhere needs the strict policy stated explicitly.
Verification as part of the play
A task after the change that reads the device and asserts the intended result turns a run from "the commands were sent" into "the outcome is correct". That is a few lines and it catches a change that applied and did something other than intended.
On a rolling change it is what makes batching meaningful: a batch whose verification fails stops the play before the next batch is touched.
# Verify the outcome, not the sending
- name: Read back what the device now has
cisco.ios.ios_l3_interfaces:
state: gathered
register: after
- name: Assert the intent was met
ansible.builtin.assert:
that:
- "'10.1.12.1/30' in (after.gathered | selectattr('name','equalto','GigabitEthernet0/1')
| map(attribute='ipv4') | first | map(attribute='address') | list)"
fail_msg: "interface address not as intended on {{ inventory_hostname }}"
Backing up first
The command pusher can retrieve the running configuration before changing anything and store it on the control machine. That gives a per-device before-state at the exact moment of the change, which is what a rollback needs.
It costs one option and a directory. Its absence is noticed only when a rollback is required, which is the worst time to discover it.
# A before-state per device, at the moment of the change
- name: Back up before touching anything
cisco.ios.ios_config:
backup: true
backup_options:
dir_path: ./backups/{{ inventory_hostname }}
filename: "{{ inventory_hostname }}-{{ ansible_date_time.iso8601_basic_short }}.cfg"
Where the controls belong
The limit and the preview are command-line habits. The batching, failure policy, verification and backup belong in the playbook, so they apply however it is run and by whoever runs it.
A playbook relying on the operator to remember flags is a playbook that will eventually be run without them.
| Control | Where | Bounds |
|---|---|---|
| Preview with difference | Command line | Everything — changes nothing |
| Limit | Command line | Which hosts |
| Batching | The playbook | How many at once |
| Failure policy | The playbook | Whether to continue |
| Verification | The playbook | Catches wrong outcomes |
| Backup | The playbook | Makes rollback possible |
Which Playbooks Cause Damage?
What are the patterns?
Four. The exact-match state value against a partial data set, which removes everything unmentioned. A run against a group that grew since the playbook was written. A change that applies and is never saved, so the devices revert on their next reload. And a template-driven configuration push where nobody read the rendered output. All four look like ordinary playbooks and all four have produced real outages.
A Deeper Dive into the Damage
The exact-match state, again
Worth repeating as the single most damaging thing available here. The value that makes a device match your data removes every object of that type your data does not list, and a data set built for three interfaces will strip the rest.
The defence is the preview, every time, without exception on any task using that value. The difference output lists the removals explicitly and they are unmistakable once seen.
The group that grew
A playbook written against a group of ten runs against the two hundred it holds now, doing exactly what it was told to a great many more devices than anybody had in mind.
Two defences. The limit on every run until the scope is deliberately confirmed, and a group whose membership is explicit rather than derived from a pattern that matches new devices automatically.
The change nobody saved
Applied to the running configuration and never written to storage. Everything works, the playbook reported success, and the devices revert at their next restart — individually, over months, as each one happens to reload.
That produces a drift pattern that is genuinely difficult to diagnose because it has no common trigger. The fix is one setting and the diagnosis, without knowing to look for it, is long.
The unreviewed template
A template rendering configuration fed straight to a device, with nobody reading the output. Everything covered elsewhere about generated text applies: a missing variable produces a line without its argument, and the first reader is the device.
Rendering to a file and diffing against the device's current configuration is the step, and it takes a minute.
# Confirm the scope before the run, every time
$ ansible-playbook site.yml --list-hosts
$ ansible-inventory --graph branch_routers
!
# And what the template will actually produce
$ ansible-playbook site.yml --check --diff --limit R1 | tee /tmp/preview.txt
A pre-run checklist
Five items. The target list is what you expect. The state values are what you intended. The preview has been read, not merely run. The first real application is limited to one device and verified. And the playbook saves what it changes.
Together they take five minutes and they cover every failure in this section.
What to keep
The preview output and the difference from each run, stored somewhere, so that a question about what a run did has an answer. The backups the run took. And a note of the scope, because "it ran against the branch group" means something different a year later.
That record is what turns a playbook run into a change with an audit trail, which is what it needs to be for anything touching production.
Blueprint framing
The CCIE Enterprise Infrastructure v1.1 blueprint includes configuration management tooling within its automation domain. What is examined is generally the interaction with the device and the behaviour of the modules rather than the tool's own internals, and the published topic list is the authority on the scope.
| Pattern | Damage | Defence |
|---|---|---|
| Exact-match state, partial data | Everything unmentioned removed | Preview, every time |
| Group grew | Far more devices than intended | List the hosts first |
| Never saved | Reverts on reload, months later | A save setting |
| Unreviewed template | Malformed configuration applied | Render and diff |
| No backup | No rollback | One option |
| No verification | Sent is not correct | An assert task |
Conclusion
The tool removes the connection handling and the mechanics, which is why it is worth using, and it hides them well enough that a task can express something far more destructive than its author intended. The five connection settings on an inventory group are what make it work at all, and their absence produces errors that read as authentication failures rather than as configuration problems.
Two kinds of module do the work. The command pusher sends lines and decides whether to bother by matching text, which is approximate in both directions — it reports changes that did not happen and misses ones that did. A resource module reads the device, computes a real difference and applies it, so its report is trustworthy and re-running is free. Prefer the second wherever one exists, and know that the state value meaning "make the device match this data" removes every object your data did not mention.
Everything that bounds a run exists because of that last sentence. Preview with the difference shown, limit to one device, verify, then batch across the rest. Put the batching, the failure policy, the backup and the verification inside the playbook rather than in a runbook, because a playbook that depends on the operator remembering flags will eventually be run without them. And list the hosts before every run — two seconds, and it is the only thing standing between a routine monthly playbook and the two hundred devices that have joined its group since it was written. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.
External Links
- RFC 4251 — The Secure Shell (SSH) Protocol Architecture
- RFC 6241 — Network Configuration Protocol (NETCONF)
- RFC 8040 — RESTCONF Protocol
- RFC 7950 — The YANG 1.1 Data Modeling Language
- RFC 3535 — Overview of the 2002 IAB Network Management Workshop
- Cisco Learning Network — CCIE Enterprise Infrastructure
Reference Notes
- RFC 4251 specifies the SSH protocol architecture, the transport used by agentless configuration management when driving network devices over their command line.
- RFC 3535 records the network management requirements identified by operators, including the need for configuration to be treated as a whole and for changes to be distinguishable from state.
- RFC 3535 identifies the value of being able to compare an intended configuration against a device's current configuration, which is the basis of declarative configuration management.
- RFC 6241 specifies NETCONF, whose candidate datastore and commit operation provide the transactional behaviour that a command-line-driven approach lacks.
- RFC 6241 distinguishes the running configuration from the startup configuration, which is why a change applied to the former does not survive a restart unless saved.
- RFC 8040 specifies RESTCONF, an alternative transport that configuration management tooling may use in place of the command line.
- RFC 7950 specifies YANG, in which the structured representations used by declarative network modules are modelled.
- RFC 7950 defines the distinction between configuration and state data, which corresponds to the read-only and configuring operations available in these tools.
- Cisco documentation describes the IOS-XE command-line configuration model, including configuration submodes, which determines the parent context a configuration line must be sent within.
- Cisco documentation states that configuration changes take effect in the running configuration immediately and are lost on reload unless copied to the startup configuration.
- Cisco documentation describes the normalisation the device applies when storing certain configuration values, which is why a textual comparison against supplied lines may not match.
- The CCIE Enterprise Infrastructure v1.1 unified exam topics include configuration management and automation tooling within the automation domain.