Calling a network device or a controller from a script is a short piece of code and a long list of things that are not obvious. The request itself is one line. Everything around it — how the credentials are presented, how long to wait, what to do when the answer is not what was expected, and whether the operation actually finished — is where the difference between a demonstration and something you would run against production lives.
Two of those are worth naming up front because they catch almost everyone. A request with no time limit waits indefinitely, and the library's default is no limit, so the first script anybody writes contains a way to hang a maintenance window on a device that accepted the connection and then stopped answering. And a controller frequently answers a change request immediately with an acknowledgement rather than a result, so a script that checks the response and stops has verified that the request was accepted and nothing else.
This article covers what a request actually needs, how the three kinds of endpoint handle credentials differently, why a script reports success when nothing changed, what makes one safe to run against production, and the failures that come from the client rather than from the device. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.
What Does a Request Actually Need?
What are the parts?
An address and a path. Headers saying what format is being sent and accepted. Credentials in whatever form the endpoint expects. A body for anything that writes. A time limit, which has no safe default and must be supplied. And a decision about certificate verification, which is on by default and which devices with self-signed certificates will fail. The last two are the ones that are never in a minimal example and always needed in real use.
A Deeper Dive into the Request
The time limit
Without one, a request waits for as long as the other end keeps the connection open without answering. A device that accepted the connection and then became busy, or whose management process stopped responding, produces exactly that condition, and the script stops there with no output and no error.
During a change window that is the worst possible failure: nothing happened, nothing was reported, and the operator does not know whether the change was applied. A limit of a few seconds for a read and rather more for a write turns it into an error that can be handled.
Certificate verification
Verification is enabled by default and a device presenting a certificate it signed itself will fail it. That is the library behaving correctly and a great deal of example code disables verification without comment, which makes it look like something you always do.
In a laboratory, disabling it is fine and should be a deliberate line rather than an inherited one. Against anything that matters, pointing the client at the certificate authority that signed the device's certificate is the correct arrangement, and it is one argument rather than a different kind of code.
Headers
Two matter for these endpoints: what format the body is in, and what format is acceptable in the answer. Getting the second wrong produces an answer in a shape the script does not expect; getting the first wrong produces a rejection.
Setting them once on a session object rather than on every request removes a whole class of inconsistency, and it is the main reason to use a session object at all beyond connection reuse.
# A session carries headers, credentials and connection reuse
import os, requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
requests.packages.urllib3.disable_warnings()
BASE = "https://10.1.1.1/restconf/data"
TIMEOUT = (5, 30) # connect, read - never omit this
s = requests.Session()
s.auth = (os.environ["NETUSER"], os.environ["NETPASS"]) # not in the file
s.headers.update({"Accept": "application/yang-data+json",
"Content-Type": "application/yang-data+json"})
s.verify = False # deliberate, lab only; use a CA bundle otherwise
Retries, and which requests may have them
A transient failure — a connection reset, a momentary unavailability — is worth retrying. A rejection is not, because retrying a bad request produces the same bad request.
The important restriction is that only operations safe to repeat may be retried automatically. A read can be retried freely. A write that creates something must not be, because a retry after a response that was lost in transit produces two of whatever was created.
# Retry transient failures on reads. Note the method list.
retry = Retry(total=3, backoff_factor=1,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET", "HEAD", "PUT", "DELETE"])
s.mount("https://", HTTPAdapter(max_retries=retry))
# POST is deliberately absent - retrying a create makes two of them.
Checking the answer before using it
A response with an error status still has a body, and that body is frequently an error description or, behind a proxy, a web page. Parsing it as though it were the expected data produces a confusing failure a long way from the cause.
Checking the status first and reading the body on failure is two lines and it converts every downstream mystery into a message naming the problem.
The body, and reading it on failure
These endpoints return structured error descriptions that name the element and the reason. That is considerably more informative than the status code and it is routinely discarded by scripts that check the code and raise a generic message.
Including the first part of the body in whatever the script reports on failure costs nothing and is the difference between "request failed" and a message that identifies which field was wrong.
| Part | Default | Safe value | Consequence of the default |
|---|---|---|---|
| Timeout | None — waits forever | Connect and read limits | Hangs a change window |
| Certificate verification | On | A CA bundle, or off knowingly | Fails against self-signed |
| Headers | Generic | Set on the session | Wrong shape or rejection |
| Retries | None | Reads only | Transient failure becomes a failure |
| Status check | None | Before using the body | Parses an error page as data |
How Are Credentials Handled?
What are the three styles?
A device's own interface takes a username and password on every request, and nothing expires. A controller typically exchanges credentials once for a token that is sent on subsequent requests and expires after a period. Some controllers use a login that establishes a session, sometimes with an additional token for write operations. The consequence that matters: a long-running script against a controller will work for a while and then fail on every request, and the failure is an expired credential rather than anything about the request.
A Deeper Dive into Credentials
The device's interface
The simplest case. Credentials accompany every request, nothing is stored, nothing expires, and a script can run for as long as it likes. The authorization requirement on the device side is the thing to have configured, and once it is, this style needs no further thought.
The credentials still need to come from somewhere other than the file, which is the same requirement everywhere and is covered below.
The token exchange
One request presenting credentials returns a token; every subsequent request presents the token in a header. The token has a lifetime, typically measured in tens of minutes, after which requests are rejected.
A script doing a single operation never notices. A script iterating over several hundred devices, or running as a scheduled job, will cross the boundary and fail partway through — having already applied part of whatever it was doing.
# Get a token, and be prepared to get another one
import base64, time, requests
def get_token(host, user, pw):
r = requests.post(f"https://{host}/dna/system/api/v1/auth/token",
auth=(user, pw), verify=False, timeout=(5, 30))
r.raise_for_status()
return r.json()["Token"]
tok = get_token(HOST, USER, PW)
s = requests.Session()
s.headers.update({"X-Auth-Token": tok, "Content-Type": "application/json"})
s.verify = False
Refreshing rather than hoping
The robust arrangement is to treat a rejection as a signal to obtain a new token and retry once. That handles expiry without needing to know the lifetime, and it handles a token invalidated for any other reason too.
Doing it by tracking the issue time and refreshing pre-emptively also works and depends on knowing the lifetime, which is documented and occasionally changes. The reactive form is more robust and is barely more code.
# Reactive refresh: one retry after an authentication rejection
def call(method, url, **kw):
global tok
kw.setdefault("timeout", (5, 30))
r = s.request(method, url, **kw)
if r.status_code in (401, 403):
tok = get_token(HOST, USER, PW)
s.headers["X-Auth-Token"] = tok
r = s.request(method, url, **kw)
return r
r = call("GET", f"https://{HOST}/dna/intent/api/v1/network-device")
r.raise_for_status()
The session style
A login establishes a session held in a cookie, and some implementations additionally require a separate token on write operations to prevent a request being made on the user's behalf by something else. Both must be present for writes to succeed, and a script that obtained only the first will read successfully and fail on every write.
That produces a distinctive symptom — reads work, writes are rejected — which points directly at the missing second credential rather than at permissions.
Where credentials come from
Not the file. Environment variables are the minimum acceptable arrangement; a secret store is better. A password in a script is a password in version control, in every copy anyone made, and in the terminal history of whoever ran it.
The two lines that read from the environment instead are the whole change, and a script with credentials in it is a finding regardless of how well the rest of it is written.
Permissions
The account a script uses should be able to do what the script does and nothing else. A script that reads inventory does not need an account that can change configuration, and giving it one means a defect in the script can do more damage than the script was ever meant to.
Separate accounts for reading and writing is a small piece of discipline with a disproportionate effect on the worst case.
| Style | Presented as | Expires | Watch for |
|---|---|---|---|
| Device interface | Username and password, each request | No | Device-side authorization |
| Token | A header, after an exchange | Yes | Long jobs failing partway |
| Session | A cookie, plus a write token | Yes | Reads work, writes rejected |
Why Does It Report Success When Nothing Changed?
What is happening?
A controller frequently answers a change request immediately, with an acknowledgement and an identifier, and performs the work afterwards. The response says the request was accepted, not that it succeeded. A script that checks the response and reports success has verified acceptance and nothing else — and the work may fail a few seconds later with nobody watching. The only honest check is to poll the identifier until the task reports finished and then read its outcome.
A Deeper Dive into Asynchronous Operations
Why controllers work this way
A change affecting several devices takes time, and holding a connection open for it is fragile. Returning an identifier immediately and letting the caller ask about progress is the sensible design, and it moves an obligation onto the caller.
That obligation is easy to miss because the immediate response looks like success. It has a success-range status code and a body, and only the presence of the identifier hints that something is outstanding.
Polling properly
Request the task's status, wait, repeat, until it reports completion or a time limit is reached. Then read the outcome, which distinguishes completed successfully from completed with an error — two different things that both count as finished.
A limit on the polling is necessary, because a task that never completes would otherwise loop indefinitely. That limit is a decision about how long the operation may reasonably take, and exceeding it is a reportable outcome rather than a reason to keep waiting.
# Accepted is not done. Poll until it is, then read the outcome.
def wait_for_task(task_id, limit=120, interval=3):
deadline = time.time() + limit
while time.time() < deadline:
r = call("GET", f"https://{HOST}/dna/intent/api/v1/task/{task_id}")
r.raise_for_status()
t = r.json()["response"]
if t.get("endTime"):
if t.get("isError"):
raise RuntimeError(f"task failed: {t.get('progress')} {t.get('failureReason')}")
return t
time.sleep(interval)
raise TimeoutError(f"task {task_id} did not finish within {limit}s")
r = call("POST", f"https://{HOST}/dna/intent/api/v1/...", json=payload)
r.raise_for_status()
wait_for_task(r.json()["response"]["taskId"])
The two kinds of finished
A task that completed and a task that completed successfully are different, and the distinction is in a field that has to be read. Treating completion as success is the second half of this mistake and it produces the same outcome for a different reason.
Both checks belong in the same helper, written once and used everywhere, so that no individual call site can omit either.
Partial success across devices
An operation spanning several devices may succeed on most and fail on some, and the overall task may report completion regardless. The per-device outcome is available and requires a further request.
Whether that matters depends on the operation. For a configuration change it matters a great deal, because a partial application across a set of devices is a state nobody designed and nobody is monitoring.
What to report
Per device: attempted, accepted, completed, and the outcome. That is four states and a script that reports only the first two is reporting that it did its part, which is not the question anybody is asking.
Printing that table at the end is a few lines and it is what makes the output trustworthy enough to act on.
# Report the outcome per device, not the acceptance
results = {}
for dev in devices:
try:
r = call("POST", f"https://{HOST}/dna/intent/api/v1/...",
json=payload_for(dev))
r.raise_for_status()
wait_for_task(r.json()["response"]["taskId"])
results[dev] = "applied"
except Exception as e:
results[dev] = f"FAILED: {e}"
for dev, outcome in sorted(results.items()):
print(f"{dev:<24} {outcome}")
Verifying the change independently
The strongest check is not the task's report but reading the device's own state afterwards and confirming it matches the intent. That catches a task that reported success and produced something other than what was asked for, which is rarer and does happen.
For a small number of devices it costs one extra request each and it converts a report into evidence. For a large number it is worth doing on a sample. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.
What Makes It Safe for Production?
What is the list?
Six properties. A time limit on every request. The status checked before the body is used. An operation that is safe to run twice. A preview mode, so the first run is not the real one. Credentials from outside the file. And an explicit list of targets rather than a query that might return more than expected. None of them is more than a few lines, and a script missing them is a demonstration rather than a tool.
A Deeper Dive into Safety
The preview mode
A flag that causes the script to report what it would do without doing it. That turns the first run into a review rather than a change, which is the difference between finding a mistake in the output and finding it in the network.
It costs one argument and one branch around the calls that write. It is the highest-value single addition to any script that changes anything, and it is absent from nearly every example.
# One flag, one branch. The first run becomes a review.
import argparse
ap = argparse.ArgumentParser()
ap.add_argument("--apply", action="store_true", help="actually make changes")
ap.add_argument("--limit", type=int, default=5, help="maximum devices to touch")
args = ap.parse_args()
for dev in devices[:args.limit]:
body = payload_for(dev)
if not args.apply:
print(f"WOULD CHANGE {dev}: {body}")
continue
r = call("POST", url_for(dev), json=body)
r.raise_for_status()
wait_for_task(r.json()["response"]["taskId"])
print(f"changed {dev}")
A limit on how far it reaches
A script that discovers its targets by querying the controller will act on whatever the query returns, which may be more than the author had in mind — particularly after the inventory grows.
Two protections. A maximum count, so an unexpectedly large result set stops rather than proceeding. And, for anything consequential, an explicit list of targets rather than a query at all. Both are a line each and both bound the worst case.
Idempotency, again
Covered elsewhere and worth restating in this context: the verb decides what a second run does. A script that is safe to re-run can be re-run after a partial failure, which is exactly when you want to re-run something.
A script that is not safe to re-run leaves an operator, after a failure halfway through, unable to do the obvious thing. That is the practical cost and it is why the property matters beyond tidiness.
Logging what happened
What was attempted, against what, with what result, timestamped. Printing to the terminal is enough for an interactive run; a scheduled one needs a file or a central destination.
The specific value is answering "what did this change" a week later. A script with no record leaves that question to a configuration diff, which works if the archive was configured and not otherwise.
# Logging that survives the terminal window
import logging
logging.basicConfig(
filename="/var/log/netchange.log",
level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s")
logging.info("start: %d devices, apply=%s", len(devices), args.apply)
# ... per device ...
logging.info("device=%s action=%s result=%s", dev, "config-push", outcome)
Concurrency, carefully
Running against devices in parallel is faster and introduces two risks: a device or controller overwhelmed by simultaneous requests, and an error that is harder to attribute because the output is interleaved.
A small number of workers — a handful rather than dozens — captures most of the speed with little of the risk. Where the endpoint publishes a rate limit, respecting it is not optional, and a rejection carrying a retry instruction should be obeyed rather than retried immediately.
Testing against something that is not production
A laboratory device, a controller's sandbox, or at worst a single production device chosen deliberately. The preview mode reduces the need for this and does not remove it, because a preview verifies the intent and not the effect.
The sequence that works: preview against everything, apply against one, verify, then apply against the rest with the limit raised. That is four steps and it is proportionate for anything touching more than a handful of devices.
| Property | Lines | Prevents |
|---|---|---|
| Timeout on every request | One argument | Hanging indefinitely |
| Status checked first | Two | Acting on an error page |
| Safe to run twice | Verb choice | Unable to re-run after a failure |
| Preview mode | One flag, one branch | The first run being the real one |
| Credentials outside the file | Two | A password in version control |
| A target limit | One | Reaching further than intended |
Which Failures Come From the Client?
What are they?
Four that are worth recognising immediately. A parse failure, because the body was not the expected format — usually an error page from something in the middle. A key error while walking the answer, because the structure is a level different from what was assumed. A certificate failure, which is the client working correctly against a self-signed device. And a rate limit rejection, which means slowing down rather than retrying. None of these is a device fault and all of them are routinely reported as one.
A Deeper Dive into Client-Side Failures
The parse failure
Attempting to interpret the body as structured data when it is not. The usual cause is a proxy, a load balancer or a portal returning a web page, which arrives with a success status and content that is not what was asked for.
Checking the content type before parsing, or simply printing the first part of the body when parsing fails, identifies it immediately. Without that, the error names a parsing problem at a character offset, which points nowhere useful.
# Parse failures: say what actually arrived
r = call("GET", url)
if r.status_code != 200:
raise SystemExit(f"{r.status_code}: {r.text[:300]}")
try:
data = r.json()
except ValueError:
raise SystemExit(f"not JSON ({r.headers.get('Content-Type')}): {r.text[:300]}")
Walking the wrong shape
Reaching into the response by name and finding nothing, because the value is inside a list rather than beside it, or because the endpoint wraps its answer in a container the example did not have.
The reliable approach is to print the whole answer once, look at it, and then write the access against what is actually there. Writing it from documentation and adjusting until it stops erroring is slower and produces code that breaks on the first response with an unusual shape.
The certificate
The client refusing a self-signed certificate is correct behaviour and reads as a failure. Disabling verification makes it go away and should be a deliberate choice recorded as such, not a reflex copied from an example.
The better arrangement for anything that matters is to give the client the certificate authority that signed the device's certificate. That is one argument and it keeps the protection that verification exists to provide.
Rate limiting
An endpoint rejecting requests because too many arrived too quickly, usually with an instruction about how long to wait. Retrying immediately makes it worse; waiting the stated period and continuing is the correct response.
On a script iterating over a large inventory this is not an edge case — it is the expected behaviour once the inventory is large enough, and handling it is what lets the script finish.
# Respect the instruction rather than retrying immediately
r = s.request(method, url, timeout=(5, 30), **kw)
if r.status_code == 429:
wait = int(r.headers.get("Retry-After", "5"))
logging.warning("rate limited, waiting %ds", wait)
time.sleep(wait)
r = s.request(method, url, timeout=(5, 30), **kw)
Paging
An endpoint returning a bounded number of results per request, with the rest available by asking again with an offset. A script that reads the first response and stops sees a portion of the inventory and reports confidently on it.
That is a quiet and consequential failure: an operation applied to "all devices" that reached the first few hundred. Checking whether the response indicates more results, and continuing until it does not, is the fix.
A diagnostic order
Status code first. If it is a failure, read the body, which usually names the problem. If it is a success and parsing failed, the body is not what was expected and printing it identifies what arrived. If parsing succeeded and a value is missing, the structure differs from the assumption and printing the whole answer settles it.
Three steps, and each one is a print statement. The instinct to reason about the code instead is what makes these take longer than they should.
Blueprint framing
The CCIE Enterprise Infrastructure v1.1 blueprint includes device programmability within its automation domain. What is examined tends to be the interaction with the device — the interface, the data, the outcome — rather than library specifics, and the published topic list is the authority on the scope.
| Failure | Origin | First move |
|---|---|---|
| Parse error | Client | Print the body and the content type |
| Missing key | Client | Print the whole response once |
| Certificate rejected | Client, correctly | A CA bundle, or disable knowingly |
| Rate limited | Endpoint, deliberately | Wait as instructed |
| Short result set | Client | Follow the pages |
| Hang with no output | Client | The missing timeout |
Conclusion
The request is one line and the parameters around it are what matter. Two of them have no safe default: there is no time limit unless you supply one, so a device that accepts a connection and stops answering hangs the script indefinitely with no output; and certificate verification is on, so a device presenting its own self-signed certificate fails until that is addressed deliberately. Both are absent from every minimal example and needed in every real one.
Credentials divide three ways. A device's interface takes them on every request and nothing expires. A controller exchanges them for a token that does, which is why a long job fails partway through having already applied half its changes — and the fix is to obtain a new token when one is rejected rather than acquiring one at the start and assuming. A session-based controller adds a second credential for writes, producing the distinctive symptom of reads working and writes being refused.
And accepted is not done. A controller answering a change request immediately has told you the request was received, not that it succeeded; the work happens afterwards and can fail with nobody watching. Poll the identifier, wait for completion, and then read whether it completed successfully, which is a different field. Add a preview flag so the first run is a review, a limit so an unexpected inventory does not become an unexpected change, and a timeout on every call — six small things, none more than a few lines, and their absence is what makes a script a demonstration. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.
External Links
- RFC 9110 — HTTP Semantics
- RFC 8040 — RESTCONF Protocol
- RFC 6749 — The OAuth 2.0 Authorization Framework
- RFC 6585 — Additional HTTP Status Codes
- RFC 8259 — The JavaScript Object Notation (JSON) Data Interchange Format
- RFC 8446 — The Transport Layer Security (TLS) Protocol Version 1.3
- Cisco Learning Network — CCIE Enterprise Infrastructure
Reference Notes
- RFC 9110 defines HTTP semantics, including which methods are idempotent and may therefore be retried safely, and which are not.
- RFC 9110 defines the 202 status code as indicating that a request has been accepted for processing without the processing having completed.
- RFC 9110 defines the 401 and 403 status codes, distinguishing a request lacking valid authentication from one whose authentication is valid but insufficient.
- RFC 6585 defines the 429 status code for rate limiting and the Retry-After header indicating how long a client should wait before repeating a request.
- RFC 8040 specifies RESTCONF, including the mapping of datastore operations onto HTTP methods and the structured error format returned in response bodies.
- RFC 8040 states that PUT replaces a target resource and PATCH merges into it, which determines whether a repeated request is safe.
- RFC 6749 describes the exchange of credentials for a time-limited access token, the pattern used by controller interfaces requiring token authentication.
- RFC 8259 specifies the JSON format, in which these interfaces exchange structured data.
- RFC 8446 specifies TLS 1.3, including certificate validation, which a client performs by default and which a self-signed certificate will not satisfy.
- Cisco documentation describes obtaining an authentication token from a controller and presenting it in a header on subsequent requests, and states that tokens have a limited lifetime.
- Cisco documentation describes asynchronous operations returning a task identifier, and the requirement to query the task resource to determine the outcome.
- The CCIE Enterprise Infrastructure v1.1 unified exam topics include device programmability within the automation domain.