A device that reacts to its own conditions without anything external being involved is the right answer to a specific and awkward problem: the thing you want to react to is frequently the network being broken, and an external system reacting to that needs the network in order to do so. A policy running on the device has no such dependency.
The mechanism is straightforward — an event of some kind occurs, a policy runs, and the policy does whatever you wrote. What is not straightforward is the small set of behaviours that are not obvious from reading a policy: the order its actions execute in is decided by how their labels compare as text, it is killed if it runs longer than a default nobody sets, and a policy that reacts to log messages can generate log messages.
This article covers what can trigger a policy, how one actually executes, when a full scripting language is the right answer instead, how to test a policy without waiting for its event, and the policies that cause the incidents they were written to detect. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.
What Can Trigger a Policy?
What are the useful sources?
Five cover almost everything. A log message matching a pattern, which is the most used and the most flexible. A timer, on a schedule or a countdown. A tracked object changing state, which connects this to every failover mechanism. A counter crossing a threshold. And a command being typed, which allows a policy to react to somebody's action. A sixth — no event at all, run only by hand — exists for testing and should be where every policy starts.
A Deeper Dive into Event Sources
The log message
A pattern matched against messages as they are generated. Since a device logs essentially everything of interest, this reaches conditions no other source names directly: an adjacency dropping, an interface state change, a specific protocol event, a threshold being crossed by something that logs it.
Its cost is that the pattern is textual and message formats differ between versions. A pattern written against one version's wording may not match another's, which makes a policy stop working after an upgrade with no indication.
! Capture state at the moment of a failure, not afterwards
R1(config)# event manager applet BGP-DOWN
R1(config-applet)# event syslog pattern "BGP-5-ADJCHANGE.*Down" maxrun 60 ratelimit 300
R1(config-applet)# action 010 cli command "enable"
R1(config-applet)# action 020 cli command "show bgp ipv4 unicast summary"
R1(config-applet)# action 030 cli command "show ip route summary"
R1(config-applet)# action 040 syslog msg "BGP down - state captured"
!
! Note the padded labels. See the next section for why.
The timer
A repeating interval, a countdown, or a calendar schedule. A repeating interval driving a check that nothing else expresses is the common use, and it pairs with a tracked object whose state the policy sets.
The thing to size carefully is the interval against the work done. A policy running every ten seconds that takes five to complete is using half the time it has, on a resource that is shared with everything the device does.
The tracked object
A state change on an object, which is what connects this to failover mechanisms. The policy runs when a circuit's usability changes, and can then do the things routing cannot — clearing state, notifying, adjusting something.
This is the highest-value pairing available, because the routing half of a failover is handled by the routing and everything else is handled here.
A command being typed
A pattern matched against commands as they are entered. That allows a policy to log an action with context, to warn, or in the extreme to prevent a command by not permitting it to proceed.
Used to block commands it is a blunt instrument with real risks, because a policy standing between an operator and the command line is a policy that can lock somebody out. Used to record, it is genuinely useful.
! React to somebody typing something
R1(config)# event manager applet LOG-WRITE-ERASE
R1(config-applet)# event cli pattern "write erase" sync yes
R1(config-applet)# action 010 syslog priority critical msg "write erase attempted by $_cli_username"
R1(config-applet)# action 020 set _exit_status 1
!
! _exit_status 0 blocks the command, 1 permits it.
! Blocking commands is a good way to lock somebody out. Be careful.
No event at all
A policy that never fires by itself and runs only when invoked. That is how a policy should be developed: written with no trigger, run by hand until it behaves, and only then given its real event.
It is also how an existing policy is tested without waiting for the condition, which is covered later and is the single most useful habit here.
Combining events
Several event statements can be tagged and combined, so that a policy runs when both of two things have happened, or either, or one within a period of the other. That expresses conditions a single event cannot.
It is worth knowing about and worth using sparingly. A policy whose trigger condition takes a paragraph to describe is a policy nobody will be able to reason about during an incident.
| Source | Reaches | Fragile because |
|---|---|---|
| Log message | Almost anything | Message wording changes between versions |
| Timer | Anything you can check | Interval against the work |
| Tracked object | Every failover mechanism | Nothing — it is a clean signal |
| Counter threshold | Error rates, utilisation | Threshold chosen without measurement |
| Command typed | Operator actions | Can block the operator |
| None | Nothing — manual only | — use it for development |
How Does an Applet Actually Run?
What decides the behaviour?
Four things that are not visible from reading the actions. The order, which is decided by comparing the labels as text rather than as numbers. The time limit, which defaults to twenty seconds and kills a policy that exceeds it partway through. The privilege the commands run at, which is not elevated unless arranged. And the variables the event made available, which are how the policy knows what happened.
A Deeper Dive into Execution
The ordering
Labels are compared as text. That means a label of ten sorts before a label of two, because the comparison is character by character and the character one precedes the character two. A policy with nine or fewer actions is unaffected; one with ten or more executes in an order its author did not write.
The usual symptom is nothing, because the actions are frequently independent. The occasional symptom is a policy that captures state after clearing it, or logs a result before computing it. Padding the labels to a fixed width removes the problem entirely.
The time limit
A policy running longer than the limit is terminated where it is. Since a policy that runs several commands against a busy device can easily take longer than the default, this produces a policy that completes its first few actions and stops.
The symptom is a partial result with no error, which reads as a policy that was written wrong. Raising the limit to cover what the policy actually does is one keyword, and it belongs on every policy that runs more than a couple of commands.
Privilege
Commands issued by a policy do not automatically run at a privileged level. A policy issuing configuration or privileged show commands without arrangement has them rejected, which appears in the policy's history as an error rather than as a permission message.
Two arrangements work: issuing the elevation command as the policy's first action, or configuring a session user for the subsystem so that its commands run as that user. The second is cleaner and interacts properly with command authorization.
! Give the subsystem an identity rather than elevating per policy
R1(config)# event manager session cli username eem-user
R1(config)# username eem-user privilege 15 algorithm-type scrypt secret <password>
!
! Or, per policy, elevate first
R1(config)# event manager applet EXAMPLE
R1(config-applet)# event none
R1(config-applet)# action 010 cli command "enable"
R1(config-applet)# action 020 cli command "show ip interface brief"
The variables
Each event source makes information available to the policy: the text of the log message that matched, the name of the interface involved, the time the event was published, the user who typed the command. Those are what let a policy act on the specifics rather than merely on the fact that something happened.
The result of the previous command is also available, which is what makes a policy able to decide. Reading a counter, comparing it, and acting on the comparison is the pattern that covers most conditional policies.
! Use what the event gave you, and what the last command returned
R1(config)# event manager applet HIGH-ERRORS
R1(config-applet)# event timer watchdog time 300 maxrun 60
R1(config-applet)# action 010 cli command "enable"
R1(config-applet)# action 020 cli command "show interfaces GigabitEthernet0/1 | include input errors"
R1(config-applet)# action 030 regexp "([0-9]+) input errors" "$_cli_result" match errs
R1(config-applet)# action 040 if $errs gt 1000
R1(config-applet)# action 050 syslog priority warnings msg "Gi0/1 input errors: $errs"
R1(config-applet)# action 060 cli command "test track 200 state down"
R1(config-applet)# action 070 else
R1(config-applet)# action 080 cli command "test track 200 state up"
R1(config-applet)# action 090 end
Rate limiting
Without one, a policy runs on every occurrence of its event. A log message appearing a hundred times a second produces a hundred policy executions a second, each issuing commands.
A minimum interval between runs bounds that. It belongs on any policy whose event can occur rapidly, which is most log-triggered policies, and its absence is how a diagnostic policy becomes the reason the device is unresponsive.
Where the output goes
Command output within a policy is captured rather than displayed, which is why a policy intended to capture state must do something with what it captured — log it, write it to a file, or include it in a message.
A policy that runs show commands and nothing else has produced no record of anything. That is a common and slightly embarrassing discovery after the incident the policy was written for.
! Capturing means writing it somewhere
R1(config)# event manager applet CAPTURE-ON-FAIL
R1(config-applet)# event syslog pattern "OSPF-5-ADJCHG.*DOWN" maxrun 90 ratelimit 600
R1(config-applet)# action 010 cli command "enable"
R1(config-applet)# action 020 file open FH flash:ospf-capture.txt a
R1(config-applet)# action 030 cli command "show ip ospf neighbor"
R1(config-applet)# action 040 file puts FH "$_cli_result"
R1(config-applet)# action 050 file close FH
R1(config-applet)# action 060 syslog msg "OSPF adjacency lost - written to flash:ospf-capture.txt"
| Behaviour | Default | Symptom when wrong |
|---|---|---|
| Action order | Text comparison of labels | Out of order above nine actions |
| Time limit | Twenty seconds | Partial execution, no error |
| Privilege | Not elevated | Commands rejected |
| Rate limit | None | Runs on every occurrence |
| Command output | Captured, not shown | Nothing recorded anywhere |
When Is a Scripting Language the Right Answer?
Where is the boundary?
Applets handle a sequence of commands with simple conditions and simple loops. Beyond that — parsing output into structure, maintaining state between runs, arithmetic of any complexity, talking to something over the network — the applet form becomes awkward and a full scripting environment becomes the sensible choice. The rough test: if the logic would be uncomfortable to explain in a sentence, it has outgrown an applet.
A Deeper Dive into the Choice
What applets do well
React, run a handful of commands, test one thing, log or write a file. That covers the overwhelming majority of what anybody wants from this, and it has a substantial advantage: the policy lives in the configuration, so it is backed up, visible in a configuration comparison and reviewable alongside everything else.
A script lives in a file on the device, which is none of those things unless somebody arranges it. That difference is worth weighing.
Where they become awkward
Parsing anything structured. Arithmetic beyond a comparison. Maintaining a value between invocations. Iterating over something derived from command output rather than over a fixed list. Each of those is possible in an applet and each produces something considerably harder to read than the equivalent script.
The signal to watch for is a policy with many conditional actions and extracted variables. At that point the applet form is being used as a programming language and it is not one.
The scripting environment
A script registered with the same subsystem, triggered by the same events, with the full facilities of a language: real data structures, real string handling, real error handling. It runs in the same context and has the same access to the device's commands.
Its cost is that it lives in a file. That file needs to be on every device, kept in step, and included in whatever backs the device up — none of which happens by itself.
! Registering a script rather than an applet
R1(config)# event manager directory user policy flash:/eem
R1(config)# event manager policy check_errors.tcl type user
!
R1# show event manager policy registered
R1# show event manager policy available
R1# dir flash:/eem
The on-device container
A third option on modern platforms: a container running a general-purpose environment, with access to the device's command line through a module. That is a considerably richer environment than either of the above and it is a different deployment problem — the container has to be enabled, resourced and maintained.
It suits something substantial running on the device. For a reaction to an event, the lighter mechanisms remain the better fit.
! The container option, for something substantial
R1(config)# iox
R1# guestshell enable
R1# guestshell run python3 /bootflash/check.py
!
# Inside, the device's CLI is available as a module
from cli import cli, configure
out = cli("show ip interface brief")
configure(["interface Loopback9", "description created from on-box python"])
Choosing between them honestly
Most reactions are a handful of commands and belong in an applet, in the configuration, where they are visible. A minority need real logic and belong in a script. A very small number need an environment and belong in a container.
The failure to avoid is reaching for the heaviest option because it is the most capable. A container maintained on three hundred branch routers to do what four applet actions would have done is a maintenance obligation nobody wanted.
Keeping scripts in step
A script on one device is a file that will diverge from the copy on another. Distributing them with the configuration management tooling, and verifying their checksums, is the arrangement that keeps them consistent.
Without it, a fleet acquires several versions of the same policy and nobody knows which device has which. That is the specific reason to prefer applets where they suffice: the configuration comparison already catches divergence. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.
| Need | Form | Lives in | Backed up by |
|---|---|---|---|
| A few commands on an event | Applet | The configuration | Everything already |
| One condition, one comparison | Applet | The configuration | Everything already |
| Parsing, state, real logic | Script | A file on the device | Nothing, unless arranged |
| Something substantial | Container | A container image | A separate process |
How Do You Test One?
What is the method?
Write it with no trigger and run it by hand. That removes the wait for the condition entirely and turns development into a loop of seconds. When it behaves, change the trigger to the real one and verify that the trigger fires by producing the condition deliberately. Then read the subsystem's own history, which records every execution with its result — the only authoritative account of what a policy actually did.
A Deeper Dive into Testing
Developing with no trigger
A policy with no event runs only when invoked, which means the edit-run-look cycle is immediate. Every action can be verified, the variables can be printed, and the ordering can be observed rather than assumed.
Doing this first is the difference between a policy that was tested and one that was reasoned about, and reasoning about a policy is exactly where the ordering and time limit behaviours catch people.
! Develop with no trigger. Run it as often as you like.
R1(config)# event manager applet UNDER-TEST
R1(config-applet)# event none
R1(config-applet)# action 010 cli command "enable"
R1(config-applet)# action 020 cli command "show ip interface brief | include Loopback"
R1(config-applet)# action 030 puts "$_cli_result"
!
R1# event manager run UNDER-TEST
Seeing the variables
Printing them during development shows exactly what the event provided and what a command returned, which is the information a policy's logic depends on. Guessing the format of a captured output and writing a pattern against the guess is how a policy silently never matches.
The print action exists for this and should be removed before the policy goes into service, because its output goes to whoever ran it and to nowhere useful when the policy fires on its own.
Triggering the real event deliberately
Once the actions work, the remaining question is whether the trigger fires. For a log-message trigger that means producing the message — shutting an interface, clearing a session, whatever generates it — and confirming the policy ran.
A pattern that does not match is the second most common fault after ordering, and it is invisible: the policy exists, is registered, and never runs. Only a deliberate trigger proves otherwise.
! Did it actually fire? The history is the authority.
R1# show event manager history events detail | begin BGP-DOWN
R1# show event manager statistics policy
R1# show event manager policy registered
!
! And while developing a pattern
R1# debug event manager action cli
R1# debug event manager detector syslog
Reading the history
The subsystem records each execution with its outcome, including whether it was terminated for exceeding the time limit. That record is the only authoritative account of what happened, and it answers the two questions that matter: did it run, and did it finish.
A policy that never appears in the history never ran, which points at the trigger. One that appears and shows a termination points at the time limit.
Testing the failure path
A policy's interesting behaviour is frequently in what it does when something is wrong — a command returns unexpected output, a file cannot be opened, a value is not what was expected. Those paths are rarely exercised in development and are exactly what runs during the incident the policy exists for.
Forcing them — by pointing the policy at something that does not exist, or by feeding it unexpected input while developing with no trigger — takes a few minutes and is the difference between a policy that helps during an incident and one that adds to it.
A pre-service checklist
Runs correctly by hand. Labels padded. Time limit covers what it does. Rate limit present if the event can repeat. Privilege arranged. Trigger produced deliberately and confirmed in the history. Output goes somewhere that survives the session. And the development print actions removed.
Eight items, a few minutes, and they cover every behaviour in this article.
! The checks, in order
R1# event manager run UNDER-TEST
R1# show running-config | section event manager applet UNDER-TEST
R1# show event manager history events detail | begin UNDER-TEST
R1# show event manager statistics policy
!
! And afterwards, confirm it is not running more than expected
R1# show event manager history events | count UNDER-TEST
Which Policies Cause Their Own Incidents?
What are the patterns?
Four. A policy triggered by log messages that itself logs, which feeds its own trigger. A policy with no rate limit reacting to an event that occurs rapidly. A policy that configures something in response to a condition its own change can recreate. And a policy that blocks commands, which can prevent the operator from removing it. Each is a loop of some kind and each has taken a device out of service.
A Deeper Dive into Self-Inflicted Problems
The logging loop
A policy matching a pattern that its own log message satisfies runs, logs, matches, runs. On a device this consumes the processor until something stops it, and the thing that would normally stop it — an operator logging in — is competing with the loop for the same resource.
Two defences. A pattern that cannot match the policy's own output, which is a matter of wording. And a rate limit, which bounds the loop's rate even if the pattern is wrong.
The rapid event
A condition that occurs many times a second — a flapping interface, a protocol failing repeatedly, a counter crossing a threshold on every sample — produces one execution per occurrence. Each execution runs commands, and commands are not free.
A rate limit is the answer and it should be the default posture rather than something added after the first incident. A policy that runs at most once a minute is almost always sufficient for a diagnostic purpose.
The self-recreating condition
A policy that changes configuration in response to a condition, where the change itself can produce that condition. Clearing a session in response to a session problem is the classic shape: the clear produces the log message that triggers the policy.
Designing this out requires thinking about what the actions produce, which is a step people skip because the actions look obviously safe individually. A rate limit also bounds it, which is the third reason to have one.
The policy that blocks the operator
A command-triggered policy that prevents commands from executing is capable of preventing the command that would remove it. That is a device requiring console access to recover, and if the console has its own restrictions, worse.
Where such a policy is genuinely wanted, it should exclude the commands needed to remove it, and the exclusion should be tested by actually using them. Reasoning that they will work is insufficient.
! If a policy blocks commands, it must not block its own removal
R1(config)# event manager applet GUARD
R1(config-applet)# event cli pattern "reload" sync yes
R1(config-applet)# action 010 syslog priority critical msg "reload attempted by $_cli_username"
R1(config-applet)# action 020 set _exit_status 1
!
! _exit_status 1 permits the command. Setting 0 blocks it -
! and a policy that blocks "no event manager applet" cannot be removed.
!
R1# show event manager policy registered
Authorization bypass
A policy can be configured so that its commands are not subject to command authorization. That exists because a policy running as a low-privilege identity cannot do much, and it means the policy's commands are not checked against the access control policy.
Anyone who can create a policy can therefore run any command through it. On a device with command authorization, that is a route around the control, and it is worth knowing about when the control is being relied upon.
What to audit
Every registered policy, what triggers it, whether it has a rate limit, whether its time limit covers what it does, and whether it bypasses authorization. That is one command per device and it finds policies added during past incidents and never removed.
Those are the ones that matter: a policy written for a condition that was resolved two years ago, still registered, still reacting to something that means nothing now.
! The audit, one command per device
R1# show event manager policy registered
R1# show running-config | section event manager applet
R1# show running-config | include event manager applet.*authorization
R1# show event manager statistics policy
!
! Policies that have never run, and ones that run constantly,
! are both worth a look.
Blueprint framing
The CCIE Enterprise Infrastructure v1.1 blueprint includes on-device event management within its automation domain. What is examined is generally the event sources, the applet structure and the interaction with other features such as object tracking, rather than extended scripting.
| Pattern | Result | Defence |
|---|---|---|
| Policy logs what triggers it | Loop, processor saturated | Rate limit and wording |
| Rapid event, no rate limit | One execution per occurrence | Rate limit by default |
| Action recreates the condition | Loop of a different shape | Consider what actions produce |
| Blocks commands | Can block its own removal | Exclude and test the escape |
| Bypasses authorization | A route around the control | Audit for it |
| Left from an old incident | Reacting to nothing | Audit registered policies |
Conclusion
The value of this is that it runs on the device, with no dependency on anything external — which matters because the condition being reacted to is frequently the network, and an external system reacting to a network failure needs the network. Triggered by a tracked object it completes a failover that routing alone leaves half done; triggered by a log message it captures state at the moment of a failure rather than twenty minutes later.
Three behaviours are not visible from reading a policy. Action labels are compared as text, so above nine actions the order runs differently from the way it was written — pad them and the problem disappears. A policy exceeding the default time limit is killed partway through with no error, which presents as a policy written wrong. And commands are not privileged unless arranged, which fails in a way that does not mention permissions.
Develop every policy with no trigger and run it by hand, because that turns a wait for a condition into a loop of seconds and it is where the ordering and time limit behaviours surface. Then trigger it deliberately and check the event history, which is the only authoritative account of whether it ran. And rate limit anything driven by log messages, always — it bounds the feedback loop, the rapid event and the self-recreating condition at once, and a diagnostic policy has never needed to run more than once a minute. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.
External Links
- RFC 5424 — The Syslog Protocol
- RFC 3164 — The BSD Syslog Protocol
- RFC 3411 — Architecture for SNMP Management Frameworks
- RFC 3877 — Alarm Management Information Base
- RFC 6192 — Protecting the Router Control Plane
- Cisco Learning Network — CCIE Enterprise Infrastructure
Reference Notes
- RFC 5424 specifies the syslog protocol, including the message structure whose text is matched by pattern-based event detection.
- RFC 5424 defines severity levels, which determine which messages a device generates and therefore which are available as triggers.
- RFC 3164 describes the earlier syslog message format still emitted by many devices, whose wording differs between implementations and versions.
- RFC 3411 describes the SNMP management framework, including the object identifiers whose values may be used as a threshold-based trigger.
- RFC 3877 describes alarm management, including the notion of an alarm being raised and cleared, which corresponds to a state-change trigger.
- RFC 6192 describes protection of the router control plane, which is the resource an unbounded event policy consumes when it executes repeatedly.
- Cisco documentation describes the Embedded Event Manager, comprising event detectors, a policy director and policies expressed as applets or as scripts.
- Cisco documentation states that applet actions are executed in ascending order of their label, with labels compared as character strings.
- Cisco documentation describes the maximum runtime for a policy, after which the policy is terminated, and states its default value.
- Cisco documentation describes the rate limit option, which specifies a minimum interval between successive executions of a policy.
- Cisco documentation describes configuring a username under which the event manager issues command-line actions, so that those commands run with appropriate privilege.
- The CCIE Enterprise Infrastructure v1.1 unified exam topics include on-device event management within the automation domain.