Modern network devices can run a container with a general-purpose operating system inside it, on the same hardware that is forwarding traffic. A script in that container has a module giving it the device's command line, so it can read state and apply configuration directly, with no session to establish and no dependency on anything being reachable.
That last property is the reason it exists. A script reacting to a network problem from outside needs the network; a script on the device does not. For a branch router that has lost its path to everything, the only thing that can still investigate is the router.
The property that deserves equal attention is the one nobody states plainly: that script can configure the device. The container is therefore a privileged execution environment with its own package manager attached, and installing a package into it is a supply chain decision about a device that carries production traffic, not a convenience. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.
What Is It For?
What problem does it solve?
Running real logic on the device, with no external dependency. Two things make that worth doing: the logic is too involved for the lightweight event mechanism, and the condition being handled is one where reaching the device from outside may not be possible. If neither applies, something simpler is the better answer, and the something simpler is usually an applet or an external script.
A Deeper Dive into the Use Case
Where an applet runs out
Parsing output into a structure. Arithmetic beyond a comparison. Keeping a value between runs. Building a request to something else. Each is possible in the lightweight form and produces something considerably harder to read than a few lines of a real language.
The signal is an applet with many extracted variables and conditional actions. At that point the applet form is being used as a programming language, and a container gives an actual one.
Where the network is the problem
A branch that has lost its path to the data centre cannot be investigated from the data centre. Anything that runs on it, runs. That is the property no external arrangement has, and it is the strongest argument for this mechanism.
The realistic use is capturing state and holding it until connectivity returns, rather than fixing anything. A script that gathers the evidence while the failure is happening is worth a great deal more than the same commands run after it has cleared.
Local processing before sending
A device producing a large amount of data — captures, counters, logs — can process it locally and send a summary rather than the raw material. On a branch with a small circuit that is a real reduction.
It is also how a device can make a decision that would otherwise require a round trip. Deciding locally and acting is faster and does not depend on the round trip succeeding.
Getting it running
Two commands: enable the hosting service, then enable the container. It takes a minute or two to start and can then be entered interactively or used to run a command.
! Two commands, then it exists
R1(config)# iox
!
R1# guestshell enable
!
R1# show iox-service
IOx Infrastructure Summary:
IOx service (CAF) : Running
IOx service (HA) : Not Supported
IOx service (IOxman) : Running
!
R1# show app-hosting list
App id State
guestshell RUNNING
Using it
Entering it gives a shell. Running a single command from the device's own prompt is the form that matters for automation, because that is what an event policy or a scheduled task invokes.
! Interactive, for development
R1# guestshell
[guestshell@guestshell ~]$ python3 --version
[guestshell@guestshell ~]$ exit
!
! One command, which is what automation uses
R1# guestshell run python3 /bootflash/check.py
R1# guestshell run df -h
!
! And the device's storage is visible from inside
R1# guestshell run ls -l /bootflash
What it is not for
Anything four applet actions would accomplish. Anything an external system can do without difficulty. And anything that would need maintaining identically across a large fleet by hand, because a container on three hundred branch routers is three hundred things that will diverge.
The honest position is that the number of problems genuinely needing this is small. When one turns up it is the right answer and nothing else is close; the rest of the time it is a heavier tool than the job requires.
| Situation | Right tool | Why |
|---|---|---|
| A few commands on an event | An applet | Lives in the configuration |
| Anything from outside | An external script | One place to maintain |
| Complex logic, on the device | The container | A real language, locally |
| Must work when the network fails | The container | No external dependency |
| Reduce data before sending | The container | Process locally, send a summary |
How Does It Reach the Network?
What is involved?
More than expected. The container has its own network stack and is not automatically connected to anything. It needs a virtual interface on the device, an address on it, a default route, and usually address translation so its traffic can leave with an address the network will route back to. Name resolution needs a file inside the container. None of it is complicated and all of it is easy to omit, which is why the first thing anybody does in there — install a package — fails.
A Deeper Dive into Container Networking
The virtual interface
A logical interface on the device forms one end of a connection to the container, and the container's own interface forms the other. Both need addresses and they need to be on the same subnet, which is a small internal range that exists only for this purpose.
That subnet is invisible to the rest of the network and needs no routing anywhere else, provided the container's outbound traffic is translated.
! The device side of the link
R1(config)# interface VirtualPortGroup0
R1(config-if)# ip address 192.168.254.1 255.255.255.0
R1(config-if)# ip nat inside
!
! The container side
R1(config)# app-hosting appid guestshell
R1(config-app-hosting)# app-vnic gateway0 virtualportgroup 0 guest-interface 0
R1(config-app-hosting-gateway0)# guest-ipaddress 192.168.254.2 netmask 255.255.255.0
R1(config-app-hosting)# app-default-gateway 192.168.254.1 guest-interface 0
Getting its traffic out
The container's address is private to the device, so its outbound traffic needs translating to something the network routes. That is ordinary address translation with the virtual interface marked as the inside and the real uplink as the outside.
Omitting it produces a container that can reach the device and nothing beyond it, which manifests as every outbound attempt timing out.
! Translate the container's traffic on the way out
R1(config)# interface GigabitEthernet0/0
R1(config-if)# ip nat outside
!
R1(config)# ip access-list standard GUESTSHELL-NAT
R1(config-std-nacl)# permit 192.168.254.0 0.0.0.255
!
R1(config)# ip nat inside source list GUESTSHELL-NAT interface GigabitEthernet0/0 overload
!
! Confirm from inside
R1# guestshell run ping -c 2 10.100.0.10
Name resolution
The container has its own resolver configuration, in its own file, and it is not populated from the device's. A container with working connectivity and no name servers fails on anything using a name, which is everything involving a package repository.
Writing the file is one line inside the container and it does not survive every recreation, so it belongs in whatever provisions the container rather than being done by hand.
Separate routing tables
Where the management path lives in its own routing table, traffic from the container needs to use it, and by default it will not. A wrapper command exists to run something in a named table, and anything needing the management path has to be invoked through it.
That is easy to forget and produces a container that works when tested against something in the ordinary table and fails against the management network.
! Run inside a named routing table, where management lives in one
R1# guestshell run sudo chvrf Mgmt-intf ping -c 2 10.200.0.50
R1# guestshell run sudo chvrf Mgmt-intf curl -s https://10.200.0.50/
!
! Which tables are available inside
R1# guestshell run sudo chvrf
Installing packages
With connectivity and resolution working, the container's package manager works normally. That is where the supply chain question arrives: a package installed there runs in an environment that can configure the device.
The defensible arrangements are an internal mirror holding reviewed packages, or vendoring the dependencies into the script's own directory so that nothing is fetched at run time. Installing from a public repository directly onto production network equipment is a decision that should be made deliberately rather than by typing the obvious command.
Whether it needs the network at all
Frequently not. A script reading device state through the command-line module and writing a file needs no network connectivity whatsoever, and that is the use case where this mechanism is strongest.
Connecting the container to the network is what makes it a general-purpose host on your network. Where the script does not need that, not doing it removes the entire section above along with the supply chain question.
| Step | Needed for | Omitting it |
|---|---|---|
| Virtual interface and addresses | Any network access | Container reaches nothing |
| Address translation | Reaching beyond the device | Everything times out |
| Name servers in the container | Anything using a name | Addresses work, names do not |
| The routing table wrapper | A separate management table | Works in tests, fails in production |
| None of it | A script using only the CLI module | Nothing — and no supply chain question |
How Does a Script Talk to the Device?
What is the mechanism?
A module available inside the container that issues commands to the device directly — no session, no credentials, no network path. One function runs a command and returns its output as text; another applies configuration from a list of lines. That directness is what makes the mechanism worth having and it is also what makes the container privileged, because anything able to import that module can configure the device.
A Deeper Dive into the Interface
Reading state
A function taking a command and returning its output. The output is the same text the command line produces, so anything that would be parsed from a terminal session is parsed the same way — which means the parsing is the work, as it always is with command output.
Where a structured alternative exists, using it instead avoids the parsing entirely. A script that reads an operational model rather than parsing a show command is shorter and does not break when the output format changes.
# Reading state, two ways
from cli import cli
# Text, which you then have to parse
out = cli("show ip interface brief")
for line in out.splitlines()[1:]:
parts = line.split()
if len(parts) >= 6 and parts[4] == "up":
print(parts[0], parts[1])
# Or read a model over the local interface and skip the parsing
# (requires container networking to the device's own address)
Applying configuration
A function taking a list of configuration lines and applying them. It enters configuration mode, sends the lines and leaves, returning whatever the device said.
Everything that applies to any configuration change applies here: check the result, make it safe to run twice, and save it afterwards if it should survive a reload. None of those happen automatically.
# Applying configuration, with the checks that are not automatic
from cli import cli, configure
result = configure([
"interface Loopback9",
"description created by on-box script",
"ip address 10.255.9.1 255.255.255.255",
])
for r in result:
if not r.success:
raise SystemExit(f"failed: {r.command} -> {r.response}")
cli("write memory") # nothing saves automatically
The privilege question
The module does not authenticate. A script that runs in the container has the device's command line, which means it can do anything the command line can do. There is no per-script permission and no separate identity.
That has two consequences. Whoever can put a file in the container can configure the device, so access to the container is equivalent to privileged access to the device. And any dependency the script pulls in runs with the same capability.
What a useful script actually looks like
Short. Read a few things, decide one thing, write the result somewhere durable. The temptation with a real language available is to build something elaborate, and an elaborate script on a network device is a maintenance obligation that will outlive whoever wrote it.
The shape that works: gather, decide, record, and let something else act on the record. A script that gathers evidence and writes it to the device's storage carries no risk of doing the wrong thing, and covers the case this mechanism exists for. It is also the version somebody else can read and modify a year later, which matters more on a device than it does anywhere else.
# Gather, decide, record. Nothing clever.
from cli import cli
import json, time
state = {
"time": time.strftime("%Y-%m-%dT%H:%M:%S"),
"ospf": cli("show ip ospf neighbor"),
"routes": cli("show ip route summary"),
"cpu": cli("show processes cpu sorted | head 5"),
}
with open("/bootflash/incident.json", "a") as f:
print(json.dumps(state), file=f)
print("captured")
Command authorization
Where the device uses per-command authorization, the question of whether the module's commands are subject to it is worth establishing rather than assuming. If they are not, the container is a route around a control the organisation is relying on.
This is the same shape as the equivalent question for the event mechanism, and it deserves the same answer: find out, document it, and decide whether it is acceptable.
Errors
A command that fails returns the device's error text rather than raising. A script that does not look will proceed as though it succeeded, which on a sequence of configuration changes applies the rest of them after the first one failed.
Checking each result is a few lines and it is the difference between a script that reports what happened and one that reports that it finished.
Being invoked
By an event policy, by a scheduled task, or by hand. The event policy form is the most useful: the lightweight mechanism detects the condition and hands off to the container for anything requiring real logic.
That division plays to both strengths — the event detection is in the configuration where it is visible, and the logic is in a language suited to it. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.
! The event mechanism detects; the container does the work
R1(config)# event manager applet ON-FAILURE
R1(config-applet)# event syslog pattern "OSPF-5-ADJCHG.*DOWN" maxrun 120 ratelimit 600
R1(config-applet)# action 010 cli command "enable"
R1(config-applet)# action 020 cli command "guestshell run python3 /bootflash/capture.py"
R1(config-applet)# action 030 syslog msg "on-box capture completed"
| Operation | Returns | Not automatic |
|---|---|---|
| Run a command | Output as text | Parsing it |
| Apply configuration | Per-line results | Checking them |
| Save | — | Doing it at all |
| Authenticate | — | There is no authentication |
What Does It Cost, and What Survives?
What are the resources?
Processor share, memory and storage, all taken from the device. The defaults are modest and can be raised, and raising them takes resources from a device whose primary job is forwarding. Storage is the one that runs out first, because a script writing captures or logs into a small container filesystem fills it quickly. As for what survives: the container and its contents survive a reload, and a software upgrade should be assumed to remove it until proven otherwise.
A Deeper Dive into the Costs
The allocation
A default profile gives a small share of each resource. That is enough for a script that reads state and writes a file and is not enough for anything doing sustained work.
The allocation can be changed, and the change requires stopping and restarting the container. Raising it substantially on a device that is busy is a decision about the device's primary function, and it is worth measuring the effect rather than assuming there is none.
! What it currently has
R1# show app-hosting resource
R1# show app-hosting detail appid guestshell | include CPU|Memory|Disk
!
! Changing it requires a stop and start
R1(config)# app-hosting appid guestshell
R1(config-app-hosting)# app-resource profile custom
R1(config-app-resource-profile-custom)# cpu 1500
R1(config-app-resource-profile-custom)# memory 512
!
R1# guestshell disable
R1# guestshell enable
Storage, which runs out
The container has a small filesystem of its own and the device's own storage is visible from inside it. A script writing output should write to the device's storage rather than into the container, both because there is more of it and because it survives the container being recreated.
A container whose filesystem fills stops working in ways that are not obviously about storage, so checking it is worth doing when something behaves oddly.
! Write to the device's storage, not into the container
R1# guestshell run df -h
Filesystem Size Used Avail Use% Mounted on
/dev/loop0 1.0G 245M 704M 26% /
/dev/sda1 11G 2.1G 8.2G 21% /bootflash
!
# In the script
with open("/bootflash/capture.txt", "a") as f:
f.write(cli("show ip ospf neighbor"))
What survives a reload
The container itself and the files in it, provided it was enabled rather than merely started. That makes it a reasonable place for a script that should be there permanently, and files on the device's own storage survive regardless.
What does not persist automatically is anything the script holds in memory, obviously, and anything installed into the container if the container is later destroyed and recreated.
What an upgrade does
Assume it removes the container. It may not, and building a process on the assumption that it will not means discovering otherwise on a fleet at an inconvenient time.
The safe arrangement is a provisioning step that recreates the container and its contents, run as part of the upgrade procedure. That also solves the divergence problem, because a recreated container is created from the same definition every time.
Keeping a fleet consistent
A container on many devices is many independent copies of whatever is inside it. Without something distributing the contents and verifying them, they diverge — different script versions, different packages, different everything.
The configuration management tooling should own this, exactly as it owns the configuration. A script distributed and verified is a script you can reason about; one placed by hand on each device is not.
Monitoring it
Whether the container is running, its resource use, and whether the expected files are present with the expected checksums. Those three answer whether the mechanism is in the state the design intended, and none of them is visible from the device's configuration.
The first is the one to alert on, because a container that stopped takes whatever depended on it with it, silently.
! The three checks, per device
R1# show app-hosting list
R1# show app-hosting resource
R1# guestshell run md5sum /bootflash/capture.py
!
! And a cheap presence check from the device itself
R1# show app-hosting detail appid guestshell | include State
| Event | Container | Files in it | Files on device storage |
|---|---|---|---|
| Reload | Survives | Survive | Survive |
| Container disabled and enabled | Survives | Survive | Survive |
| Container destroyed | Gone | Gone | Survive |
| Software upgrade | Assume gone | Assume gone | Usually survive |
Which Deployments Are Not Worth It?
What should be avoided?
Four. Using it for something a handful of applet actions would do, which trades a configuration-resident mechanism for a file-based one for no gain. Using it where an external system would work, which puts maintenance on every device instead of one. Installing packages from a public source onto production devices without a review path. And deploying it across a large fleet without a provisioning mechanism, which produces divergence that nobody is tracking.
A Deeper Dive into the Wrong Deployments
Using it instead of an applet
An applet lives in the configuration, so it is backed up, appears in a configuration comparison, and is reviewed alongside everything else. A script in a container is a file that none of those apply to.
Where the work fits in a few actions, the applet is better on every axis except expressiveness. Reaching for the container because it is more capable is choosing a maintenance obligation over a configuration line.
Using it instead of an external system
Anything that can be done from outside should be, because there is one copy to maintain rather than one per device. The exception is the case that motivates this whole mechanism — work that must happen when the device is unreachable — and that exception covers less than people expect.
A useful test: if the script would work perfectly well run from a server, run it from a server.
The package question
A package installed into the container runs with the ability to configure the device. Pulling one from a public repository onto a production router is introducing code from outside the organisation into a privileged position on network infrastructure.
That may be acceptable, and it should be a decision rather than a side effect of typing the obvious command. The arrangements that make it defensible are an internal mirror of reviewed packages, or vendoring dependencies into the script so nothing is fetched at run time.
# Vendor the dependencies rather than fetching at run time
# On a build machine:
$ pip3 install --target ./vendor requests
$ tar czf payload.tgz script.py vendor/
!
# On the device: copy it in, extract, and point at it
R1# copy tftp://10.200.0.70/payload.tgz bootflash:
R1# guestshell run tar xzf /bootflash/payload.tgz -C /bootflash/app
!
# In the script:
import sys; sys.path.insert(0, "/bootflash/app/vendor")
import requests
Fleet deployment without provisioning
Placing a script by hand on ten devices is manageable and produces ten copies that will diverge. On three hundred it is not manageable at all.
The provisioning has to exist before the deployment, not after, and it has to include recreating the container after an upgrade. Without both, the deployment degrades quietly over the following year.
A decision checklist
Would an applet do it? Would an external script do it? Does it need the network from inside the container, and if so is the supply chain question answered? Is there a provisioning mechanism? And is the recreation step in the upgrade procedure?
Five questions. Two yes answers to the first two mean this is the wrong tool, and a no to any of the last three means the deployment is not ready.
! Before deploying to more than one device
R1# show app-hosting list
R1# show app-hosting detail appid guestshell | include State|Memory|CPU
R1# guestshell run md5sum /bootflash/app/script.py
R1# show running-config | include ^iox|app-hosting|VirtualPortGroup
!
! And confirm the provisioning reproduces all of it from nothing
R1# guestshell destroy
! ... run the provisioning ... then verify the checks above again
Removing it
A container enabled for a project and left behind consumes resources and holds code nobody is maintaining. Destroying it when the project ends is one command and it is routinely skipped.
An audit for enabled containers across the estate finds them, and the ones it finds are usually from something that concluded a year earlier.
The honest summary
This is the right answer to a narrow problem: real logic that must run on the device because the device may be all there is. For that problem nothing else comes close. For everything else, something lighter exists and is better on maintenance, visibility and review — and the temptation to use the most capable tool available is the main way this ends up somewhere it does not belong.
Blueprint framing
The CCIE Enterprise Infrastructure v1.1 blueprint includes on-device programmability within its automation domain. What is examined is generally what the environment provides and how a script reaches the device, rather than extended container administration.
| Deployment | Problem | Better |
|---|---|---|
| Replaces a few applet actions | File instead of configuration | An applet |
| Could run externally | Maintenance per device | An external script |
| Public packages on production | Unreviewed code, privileged | A mirror, or vendoring |
| Fleet, no provisioning | Silent divergence | Provision before deploying |
| No upgrade step | Stops after each upgrade | Recreation in the procedure |
| Left behind after a project | Unmaintained privileged code | Destroy it |
Conclusion
The mechanism runs a general-purpose environment on the device, with a module giving a script the device's command line directly. Its value is a narrow one and it is genuine: work that must happen when the network is the problem has nowhere else to run. For capturing state during a failure, on a branch that has lost its path to everything, nothing else is available at all.
The property to state plainly is that the module does not authenticate. Any code in the container can configure the device, which makes the ability to place a file there equivalent to privileged access, and makes installing a package a decision about introducing outside code into a privileged position on production infrastructure. Vendoring dependencies or mirroring reviewed packages are the arrangements that make it defensible; typing the obvious command is not.
And the honest scope is small. An applet lives in the configuration where it is backed up and reviewed; an external script has one copy to maintain instead of one per device. Both are better on every axis except expressiveness. Before deploying this, answer whether either would do the job — and if the deployment goes ahead, provision the container from a single definition and put its recreation in the upgrade procedure, because a software upgrade should be assumed to remove it and a fleet of hand-placed copies diverges within a year. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.
External Links
- RFC 6192 — Protecting the Router Control Plane
- RFC 3022 — Traditional IP Network Address Translator
- RFC 1035 — Domain Names: Implementation and Specification
- RFC 8040 — RESTCONF Protocol
- RFC 4949 — Internet Security Glossary
- Cisco Learning Network — CCIE Enterprise Infrastructure
Reference Notes
- RFC 6192 describes the router control plane and the resources consumed by processes running on the device, which a hosted application shares.
- RFC 3022 specifies address translation, the mechanism by which a hosted application's private address is translated for traffic leaving the device.
- RFC 1035 specifies the domain name system, whose resolver configuration a hosted environment maintains separately from the host device's own settings.
- RFC 8040 specifies RESTCONF, which a script may use to read structured data rather than parsing command output.
- RFC 4949 defines least privilege, the principle against which an execution environment with unrestricted command-line capability should be assessed.
- Cisco documentation describes the application hosting infrastructure and the guest shell as a container running a Linux environment on the device.
- Cisco documentation describes enabling application hosting and the guest shell, and the commands for entering it or running a single command within it.
- Cisco documentation describes the virtual port group and application virtual interface configuration required for the container to have network connectivity.
- Cisco documentation describes the Python module available inside the guest shell for issuing exec and configuration commands to the host device.
- Cisco documentation describes the application resource profile controlling the processor, memory and storage available to a hosted application.
- Cisco documentation notes that the guest shell persists across reloads and that a software upgrade may remove or reset it.
- The CCIE Enterprise Infrastructure v1.1 unified exam topics include on-device programmability within the automation domain.