Latest Cisco, PMP, AWS, CompTIA, Microsoft Materials on SALE Get Now Get Now

Ansible for IOS-XE: The Task Configured Three Interfaces and Cleared Forty-Five

Driving network devices from a playbook is popular because it removes the two things people dislike about scripting: there is no connection handling to write, and the description of what you want is separate from the mechanics of getting there. Both of those are real and both make the tool worth using.

What they also do is hide the mechanics well enough that a playbook can express something considerably more destructive than its author intended. One word in a task decides whether the tool adds your configuration to what is there or makes the device match your data and removes everything else. Both spellings look equally routine, and the second one run against a partial data set is a configuration eraser.

This article covers how the tool reaches a device at all, the difference between the module that pushes commands and the modules that describe state, what the word "changed" actually means in the output, how to bound what a run can affect, and the playbooks that cause damage. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.

Blog ClaimOne state value means "make the device match this data and remove everything else" — so a resource module run against a partial data set is a configuration eraser with a reassuring name, and it looks identical to the safe spelling.
The state value decides whether a task adds to the device or makes the device match your data, and everything that bounds a run — preview, limit, batching — exists because that distinction is a single word.

How Does It Reach a Device?

What has to be right?

An inventory entry naming the device, a connection type suited to network equipment rather than to servers, the platform's name so the right module implementations are used, credentials, and a statement that privileged mode is required. Getting the connection type or the platform name wrong produces a failure that reads like an authentication problem, because the tool tries to behave as though the device were a server and the device does not cooperate.

A Deeper Dive into Connectivity

Why the connection type matters

The default behaviour assumes a general-purpose operating system with a shell, a file system it can copy a module into, and an interpreter to run it. A network device has none of those in the expected form, so the tool has a separate connection type that keeps the logic on the control machine and sends commands over the session.

Omitting it produces a failure about being unable to execute something on the device, which reads as a permissions problem and is a connection type problem.

# Inventory: the four settings that make it work at all
all:
  children:
    ios:
      hosts:
        R1: {ansible_host: 10.1.1.1}
        R2: {ansible_host: 10.1.1.2}
      vars:
        ansible_connection: ansible.netcommon.network_cli
        ansible_network_os: cisco.ios.ios
        ansible_user: "{{ lookup('env', 'NETUSER') }}"
        ansible_password: "{{ lookup('env', 'NETPASS') }}"
        ansible_become: true
        ansible_become_method: enable
        ansible_command_timeout: 60

Privileged mode

Configuration requires it and the session does not start there unless told. A playbook without the escalation settings connects successfully, runs a read-only task fine, and fails on anything that configures — which makes it look like the configuration task is wrong rather than the connection.

Where the account already lands in privileged mode, the setting is harmless. Including it unconditionally is simpler than deciding per device.

Fact gathering

The default behaviour of collecting information about the host assumes a server and either fails or wastes time on a network device. Turning it off in the play and using the platform's own information-gathering module where facts are needed is the correct arrangement.

This is one line and its absence is a common cause of a playbook that takes far longer than it should, or that fails at the very first step for reasons unrelated to the task.

# The play header for network devices
- name: Configure interfaces
  hosts: ios
  gather_facts: false
  connection: ansible.netcommon.network_cli
  tasks:

    - name: Collect device facts properly
      cisco.ios.ios_facts:
        gather_subset: [hardware, interfaces]

    - name: Show what we learned
      ansible.builtin.debug:
        msg: "{{ ansible_net_version }} on {{ ansible_net_model }}"

Credentials

Not in the inventory file as plain text. Either from the environment, or from the tool's own encrypted storage, or from a secret manager. The inventory is committed to version control and a password in it is a password in version control.

The encrypted storage is the lowest-friction option that is actually acceptable, and it integrates with the inventory directly.

Timeouts

A device that is slow to answer — a large configuration, a busy processor — will exceed the default command timeout and produce a failure that looks like a connectivity problem. Raising it for playbooks that read large outputs is routine.

The persistent connection also has its own idle timeout, and a play with a long gap between tasks against the same host can lose its session. Both are settings rather than problems once you know they exist.

Pitfall: every task fails with what looks like an authentication error Symptom: a playbook connects and every task fails, with messages suggesting the module could not be executed or the credentials were rejected. The same credentials work perfectly for an interactive session to the same device. Cause: the connection type was not set for network equipment, so the tool attempted to copy and execute a module on the device as though it were a server. Confirm: the inventory or play has no network connection type set, or names a general-purpose one. Fix: set the network connection type and the platform name in the inventory group, which is two lines and applies to every host in it.
Setting Omitting it produces
Network connection type Failures that look like authentication
Platform name Wrong module implementation, odd errors
Privilege escalation Reads work, configuration fails
Fact gathering disabled Slow starts or an immediate failure
Command timeout raised Large outputs fail as connectivity problems
Put the connection settings on the groupFive settings on an inventory group cover every host in it, and every one of them is the kind of thing that produces a misleading error when absent. Getting the group right once removes them from every playbook that ever runs against it.
Sub claimA missing connection type produces a failure that reads as an authentication problem, which is why the first debugging move on a new inventory should be the group variables rather than the credentials.

Command Pusher or Resource Module?

What is the difference?

The command pusher takes lines of configuration and sends them, deciding whether to bother by matching the text against the running configuration. That match is approximate, so it sometimes reports a change that did not happen and sometimes misses one that should. A resource module takes a description of the desired state as structured data, reads what the device currently has, computes the difference and sends only that. The second is genuinely idempotent; the first is idempotent-ish.

A Deeper Dive into the Two Kinds

What the command pusher actually does

It checks whether the lines you supplied appear in the running configuration within the context you named, and if they do not, it enters that context and sends them. The check is textual, so a line that is present in a slightly different form — a different abbreviation, a value the device normalised — is not recognised.

The result is a task that reports a change on every run without changing anything, which is noise, or a task that reports no change when the configuration has drifted, which is worse.

# The command pusher: you supply lines, it matches text
- name: Configure an interface description
  cisco.ios.ios_config:
    parents: "interface GigabitEthernet0/1"
    lines:
      - "description Uplink to core"
      - "ip address 10.1.12.1 255.255.255.252"
    match: line
    save_when: modified
!
# Works anywhere, for anything. Idempotency is textual and approximate.

What a resource module does

It reads the relevant part of the configuration, parses it into structured data, compares that against the data you supplied, computes the commands needed to reconcile the two, and sends them. If nothing is needed, it sends nothing and reports no change.

That is a real comparison rather than a text match, so the report is trustworthy and re-running is genuinely free. Its limitation is coverage: it handles what the module models and nothing else.

# A resource module: you describe the state, it works out the commands
- name: Ensure interface configuration matches
  cisco.ios.ios_l3_interfaces:
    config:
      - name: GigabitEthernet0/1
        ipv4:
          - address: 10.1.12.1/30
      - name: Loopback0
        ipv4:
          - address: 10.255.0.1/32
    state: merged

When to use which

A resource module wherever one covers what you are configuring, because the idempotency is real and the data is easier to review than a list of commands. The command pusher for anything not covered, which on a large platform is a lot of things.

Mixing both in one playbook is normal and fine. What is worth avoiding is using the command pusher for something a resource module covers, because that gives up the guarantee for no benefit.

The state values, and the one to be careful with

Adding your data to what is there. Rewriting the objects you named while leaving others alone. Making the device match your data entirely, which removes every object of that kind that your data does not mention. Removing what you named. And three read-only values that change nothing.

The third one is the dangerous one and it is spelled almost the same as the second. Run against a data set listing three interfaces on a switch with forty-eight, it removes the configuration from forty-five of them.

Pitfall: a task removes configuration from every object it did not mention Symptom: a playbook intended to configure three interfaces is run and every other interface on the device loses its configuration. The task reported success and its data listed exactly the three interfaces intended. Cause: the state value instructs the module to make the device match the supplied data exactly, which means removing every object of that type not present in it. The data was a partial list and the module treated it as the complete desired state. Confirm: the task's state value is the one meaning "make the device match this"; the diff shows removals for objects never mentioned. Fix: use the state value that merges, unless the data genuinely is the complete intended configuration for that resource on that device — and preview with the check flag before any run that uses the exact-match value.

The read-only states

One reads the device and returns the structured data, which is how you discover what the module's data model looks like for something already configured. One takes your data and produces the commands it would send, without touching a device at all. One takes a configuration text and parses it into the module's data model.

The second is the most useful and the least used. It converts a playbook into a review artefact in one run with no device involved, which is the cheapest possible safety step.

# Produce the CLI without touching anything
- name: Show what would be sent
  cisco.ios.ios_l3_interfaces:
    config: "{{ interface_config }}"
    state: rendered
  register: preview

- ansible.builtin.debug:
    var: preview.rendered
!
# And learn the data model from a configured device
- name: Read the current state as data
  cisco.ios.ios_l3_interfaces:
    state: gathered
  register: current

Templates and the command pusher

A template rendering a block of configuration, fed to the command pusher, is a common and workable pattern for things no resource module covers. The template's output is text and inherits every property of generated text, including that nothing validates it.

Where that pattern is used, rendering to a file and reviewing it before the run is the safety step, exactly as it is for any generated configuration.

Need Use Why
Something a module covers Resource module Real idempotency, reviewable data
Anything else Command pusher Covers everything, approximately
Read state as data gathered Learn the model from a real device
See the commands first rendered No device involved at all
Read-only checks The command runner Never reports a change
Sub claimThe state value meaning "make the device match this" is spelled almost identically to the one meaning "rewrite what I named", and the difference between them on a partial data set is the rest of the device's configuration.

What Does "Changed" Actually Mean?

Can it be trusted?

It depends entirely on the module. From a resource module it means a real difference was computed and applied, and it is trustworthy. From the command pusher it means a text match failed, which may or may not correspond to an actual difference. From the read-only command runner it never appears at all, regardless of what the commands returned. A summary line counting changes therefore means different things for different tasks in the same play.

A Deeper Dive into the Report

The perpetually changing task

A command pusher task whose lines never match the device's rendering of them reports a change on every run and sends the same commands every time. The configuration is correct throughout; the report is wrong.

Common causes: a value the device normalises to a different form, a command that appears in the configuration under a different parent, or an abbreviation. The fix is to write the lines exactly as the device renders them, which means reading them from the device rather than from documentation.

Pitfall: a task reports a change on every run and nothing is ever different Symptom: a playbook run repeatedly against an unchanged device reports the same task as changed every time. Inspecting the device shows the configuration is already exactly as intended. Cause: the module decides by matching the supplied text against the running configuration, and the device renders that configuration in a slightly different form — a normalised value, a different abbreviation, a different parent context. The match fails and the commands are sent again. Confirm: compare the supplied lines character by character with how the device displays them. Fix: write the lines exactly as the device renders them, taken from the device itself, or move to a resource module where one exists and the comparison is structural rather than textual.

The task that misses a difference

The mirror image, and the more serious one. A device whose configuration has drifted in a way the text match does not notice reports no change, so the playbook run confirms compliance that does not exist.

This is the argument for resource modules stated most sharply: a compliance run built on approximate matching reports compliance approximately, which for a compliance run is the one property it must not have.

Reading the difference rather than the summary

The diff output shows what was actually sent, per device, which is the only trustworthy view. A summary line saying three hosts changed says nothing about what changed on them.

Running with the difference displayed, always, converts the output from a count into a record. It is one flag and it makes the run reviewable.

# Always. The summary line is not enough.
$ ansible-playbook interfaces.yml --check --diff --limit R1
!
$ ansible-playbook interfaces.yml --diff --limit R1
!
# The diff shows the commands per host, which is the record.

Check mode

The tool can run a play and report what it would do without doing it. For resource modules that is accurate, because the difference is computed the same way either way. For the command pusher it is approximate for the same reason its change reporting is.

It is still worth using on both, because approximate is a great deal better than nothing when the alternative is applying a change blind.

Saving the configuration

A change applied to the running configuration is not saved unless something saves it. The command pusher has a setting for this and it is not on by default; resource modules generally do not save either.

A playbook that changes configuration and never saves produces devices that revert on reload, which is discovered weeks later. The save belongs either as a setting on the tasks that change things or as a final task in the play.

# Nothing saves automatically. Two ways to do it.
- name: Configure, and save only if something changed
  cisco.ios.ios_config:
    lines: ["description managed by automation"]
    parents: "interface Loopback0"
    save_when: modified
!
# Or once, at the end of the play
- name: Save if anything changed
  cisco.ios.ios_config:
    save_when: always
  when: ansible_play_hosts_all | length > 0

Handlers

A task that should run only if something else changed — saving, or restarting a process — is expressed as a handler triggered by the change notification. That is the idiomatic arrangement and it depends on the change reporting being accurate.

With a task that reports a change every run, the handler fires every run. That turns an inaccurate report into an action, which is how a noisy playbook becomes a playbook that writes to storage on every scheduled execution. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.

Module changed means Trust the report
Command runner Never set n/a — read only
Command pusher A text match failed Approximately
Resource module A computed difference was applied Yes
In check mode Would have done the above Same caveats
Run with the difference displayed, alwaysThe summary counts hosts; the difference shows commands. On a change that matters, the second is the record of what happened and the first is a number. It is one flag and there is no reason ever to omit it.
Sub claimA compliance run built on approximate text matching reports compliance approximately, which is the one property a compliance run must not have.

How Do You Bound What a Run Affects?

What are the controls?

Four, and they compose. Preview mode with the difference shown, which changes nothing. A limit restricting the run to named hosts, so the first real run touches one device. Batching, so a run across many devices proceeds a few at a time and stops if a batch fails. And a failure policy deciding whether one host's failure stops the play. Used together they turn a playbook run from an event into a procedure.

A Deeper Dive into Blast Radius

The limit

Restricting a run to named hosts or a pattern. The first application of any new or modified playbook should use it with a single host, and the second with a handful.

It is also the protection against an inventory that grew. A playbook written when the group held ten devices runs against the two hundred it holds now, and the limit is what makes that a deliberate decision rather than a discovery.

# The sequence, four commands
$ ansible-playbook site.yml --check --diff --limit R1     # see it
$ ansible-playbook site.yml --diff --limit R1             # do one
$ ansible-playbook site.yml --diff --limit 'branch[0:4]'  # do a few
$ ansible-playbook site.yml --diff                        # the rest

Batching

By default the tool works on many hosts in parallel and completes each task across all of them before moving to the next. Batching changes that to completing the whole play on a subset before starting the next subset.

For a configuration change that matters, batching is what limits a bad change to one batch. Combined with a failure policy that stops the play when a batch fails badly enough, it means a mistake reaches a fraction of the estate rather than all of it.

# A few at a time, and stop if a batch goes wrong
- name: Roll out the change
  hosts: branch_routers
  gather_facts: false
  serial: 5
  max_fail_percentage: 20
  any_errors_fatal: false
  tasks:
    - name: Apply
      cisco.ios.ios_config:
        src: templates/branch.j2
        save_when: modified

The failure policy

By default a host that fails is removed from the play and the others continue. For an independent change that is right; for a change where a partial application across the estate is worse than none, stopping everything on the first failure is right.

Deciding which applies is a property of the change rather than a default to accept. A change that must be everywhere or nowhere needs the strict policy stated explicitly.

Verification as part of the play

A task after the change that reads the device and asserts the intended result turns a run from "the commands were sent" into "the outcome is correct". That is a few lines and it catches a change that applied and did something other than intended.

On a rolling change it is what makes batching meaningful: a batch whose verification fails stops the play before the next batch is touched.

# Verify the outcome, not the sending
- name: Read back what the device now has
  cisco.ios.ios_l3_interfaces:
    state: gathered
  register: after

- name: Assert the intent was met
  ansible.builtin.assert:
    that:
      - "'10.1.12.1/30' in (after.gathered | selectattr('name','equalto','GigabitEthernet0/1')
                            | map(attribute='ipv4') | first | map(attribute='address') | list)"
    fail_msg: "interface address not as intended on {{ inventory_hostname }}"

Backing up first

The command pusher can retrieve the running configuration before changing anything and store it on the control machine. That gives a per-device before-state at the exact moment of the change, which is what a rollback needs.

It costs one option and a directory. Its absence is noticed only when a rollback is required, which is the worst time to discover it.

# A before-state per device, at the moment of the change
- name: Back up before touching anything
  cisco.ios.ios_config:
    backup: true
    backup_options:
      dir_path: ./backups/{{ inventory_hostname }}
      filename: "{{ inventory_hostname }}-{{ ansible_date_time.iso8601_basic_short }}.cfg"

Where the controls belong

The limit and the preview are command-line habits. The batching, failure policy, verification and backup belong in the playbook, so they apply however it is run and by whoever runs it.

A playbook relying on the operator to remember flags is a playbook that will eventually be run without them.

Control Where Bounds
Preview with difference Command line Everything — changes nothing
Limit Command line Which hosts
Batching The playbook How many at once
Failure policy The playbook Whether to continue
Verification The playbook Catches wrong outcomes
Backup The playbook Makes rollback possible
Sub claimA playbook relying on the operator to remember command-line flags will eventually be run without them, which is why batching, failure policy and backup belong inside it rather than in a runbook.

Which Playbooks Cause Damage?

What are the patterns?

Four. The exact-match state value against a partial data set, which removes everything unmentioned. A run against a group that grew since the playbook was written. A change that applies and is never saved, so the devices revert on their next reload. And a template-driven configuration push where nobody read the rendered output. All four look like ordinary playbooks and all four have produced real outages.

A Deeper Dive into the Damage

The exact-match state, again

Worth repeating as the single most damaging thing available here. The value that makes a device match your data removes every object of that type your data does not list, and a data set built for three interfaces will strip the rest.

The defence is the preview, every time, without exception on any task using that value. The difference output lists the removals explicitly and they are unmistakable once seen.

The group that grew

A playbook written against a group of ten runs against the two hundred it holds now, doing exactly what it was told to a great many more devices than anybody had in mind.

Two defences. The limit on every run until the scope is deliberately confirmed, and a group whose membership is explicit rather than derived from a pattern that matches new devices automatically.

Pitfall: a routine playbook reaches ten times the devices it used to Symptom: a playbook that has been run monthly without incident affects a large number of devices, including ones nobody associated with it. Nothing about the playbook changed. Cause: its target group is defined by a pattern, and devices added to the inventory since it was written now match that pattern. The playbook did exactly what it was told to everything that qualified. Confirm: list the hosts the playbook would target and compare against the number anybody expects. Fix: confirm the target list before every run with the host-listing option, and prefer explicit group membership over patterns for anything that changes configuration.

The change nobody saved

Applied to the running configuration and never written to storage. Everything works, the playbook reported success, and the devices revert at their next restart — individually, over months, as each one happens to reload.

That produces a drift pattern that is genuinely difficult to diagnose because it has no common trigger. The fix is one setting and the diagnosis, without knowing to look for it, is long.

The unreviewed template

A template rendering configuration fed straight to a device, with nobody reading the output. Everything covered elsewhere about generated text applies: a missing variable produces a line without its argument, and the first reader is the device.

Rendering to a file and diffing against the device's current configuration is the step, and it takes a minute.

# Confirm the scope before the run, every time
$ ansible-playbook site.yml --list-hosts
$ ansible-inventory --graph branch_routers
!
# And what the template will actually produce
$ ansible-playbook site.yml --check --diff --limit R1 | tee /tmp/preview.txt

A pre-run checklist

Five items. The target list is what you expect. The state values are what you intended. The preview has been read, not merely run. The first real application is limited to one device and verified. And the playbook saves what it changes.

Together they take five minutes and they cover every failure in this section.

What to keep

The preview output and the difference from each run, stored somewhere, so that a question about what a run did has an answer. The backups the run took. And a note of the scope, because "it ran against the branch group" means something different a year later.

That record is what turns a playbook run into a change with an audit trail, which is what it needs to be for anything touching production.

Blueprint framing

The CCIE Enterprise Infrastructure v1.1 blueprint includes configuration management tooling within its automation domain. What is examined is generally the interaction with the device and the behaviour of the modules rather than the tool's own internals, and the published topic list is the authority on the scope.

Pattern Damage Defence
Exact-match state, partial data Everything unmentioned removed Preview, every time
Group grew Far more devices than intended List the hosts first
Never saved Reverts on reload, months later A save setting
Unreviewed template Malformed configuration applied Render and diff
No backup No rollback One option
No verification Sent is not correct An assert task
List the hosts before every runOne command shows exactly which devices a playbook will target. It takes two seconds, it is the only defence against an inventory that grew, and the failure it prevents is the one that affects the most devices.
Sub claimA change applied and never saved produces devices reverting individually over months as each one happens to reload, which is a drift pattern with no common trigger and therefore no obvious cause.

Conclusion

The tool removes the connection handling and the mechanics, which is why it is worth using, and it hides them well enough that a task can express something far more destructive than its author intended. The five connection settings on an inventory group are what make it work at all, and their absence produces errors that read as authentication failures rather than as configuration problems.

Two kinds of module do the work. The command pusher sends lines and decides whether to bother by matching text, which is approximate in both directions — it reports changes that did not happen and misses ones that did. A resource module reads the device, computes a real difference and applies it, so its report is trustworthy and re-running is free. Prefer the second wherever one exists, and know that the state value meaning "make the device match this data" removes every object your data did not mention.

Everything that bounds a run exists because of that last sentence. Preview with the difference shown, limit to one device, verify, then batch across the rest. Put the batching, the failure policy, the backup and the verification inside the playbook rather than in a runbook, because a playbook that depends on the operator remembering flags will eventually be run without them. And list the hosts before every run — two seconds, and it is the only thing standing between a routine monthly playbook and the two hundred devices that have joined its group since it was written. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.

Reference Notes

  1. RFC 4251 specifies the SSH protocol architecture, the transport used by agentless configuration management when driving network devices over their command line.
  2. RFC 3535 records the network management requirements identified by operators, including the need for configuration to be treated as a whole and for changes to be distinguishable from state.
  3. RFC 3535 identifies the value of being able to compare an intended configuration against a device's current configuration, which is the basis of declarative configuration management.
  4. RFC 6241 specifies NETCONF, whose candidate datastore and commit operation provide the transactional behaviour that a command-line-driven approach lacks.
  5. RFC 6241 distinguishes the running configuration from the startup configuration, which is why a change applied to the former does not survive a restart unless saved.
  6. RFC 8040 specifies RESTCONF, an alternative transport that configuration management tooling may use in place of the command line.
  7. RFC 7950 specifies YANG, in which the structured representations used by declarative network modules are modelled.
  8. RFC 7950 defines the distinction between configuration and state data, which corresponds to the read-only and configuring operations available in these tools.
  9. Cisco documentation describes the IOS-XE command-line configuration model, including configuration submodes, which determines the parent context a configuration line must be sent within.
  10. Cisco documentation states that configuration changes take effect in the running configuration immediately and are lost on reload unless copied to the startup configuration.
  11. Cisco documentation describes the normalisation the device applies when storing certain configuration values, which is why a textual comparison against supplied lines may not match.
  12. The CCIE Enterprise Infrastructure v1.1 unified exam topics include configuration management and automation tooling within the automation domain.