Latest Cisco, PMP, AWS, CompTIA, Microsoft Materials on SALE Get Now Get Now

Model-Driven Telemetry: You Configured Ten Seconds and Got Ten a Second

Polling a network for its state is a design that stopped scaling some time ago. A collector asking a thousand devices for the same values every thirty seconds is thirty thousand request-and-response exchanges a minute, timed by the collector's clock, each one a small amount of work on a device that had better things to do. The state it produces is a series of snapshots the collector happened to take, not a record of what the device experienced.

Subscribing inverts it. The device is told once which part of its model to report and how often, and it pushes updates with its own timestamps for as long as the subscription exists. No polling loop, no per-request overhead, and a record timed by the thing being measured rather than by the thing measuring it.

This article covers what is actually being subscribed to, the choice between the device connecting outward and a collector connecting inward, what the interval means — which is not what most people assume — what reporting only on change requires, and why a correctly configured subscription delivers nothing. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.

Blog ClaimThe periodic interval is expressed in hundredths of a second, so a subscription configured with what looks like a ten-second interval is sending ten times a second — and the first sign of it is the collector falling over rather than anything on the device.
A subscription replaces a polling loop with a push timed by the device's own clock, and the three settings that decide whether it works are the filter, the interval unit and which end opens the connection.

What Is Being Subscribed To?

What does a subscription name?

A path into the device's model, exactly as a configuration request does — the same models, the same tree, the same addressing. The difference is that a configuration request reads a value once and a subscription asks for it repeatedly without being asked again. So the first thing a subscription needs is a path that exists on that device, and the most common reason one never delivers anything is that the path names something the device's software version does not implement.

A Deeper Dive into the Subscription

The same models as everything else

Whatever describes the device's operational state to a configuration protocol describes it here. That means the path can be discovered the same way — ask the device what it implements, navigate down to the thing you want — and that a path verified with a single read is a path that will work in a subscription.

That is the efficient development sequence: read the value once over an ordinary interface, confirm it is what you want, then subscribe to the path that returned it.

# Verify the path with a single read before subscribing to it
$ curl -sk -u admin:pass -H "Accept: application/yang-data+json"    https://10.1.1.1/restconf/data/Cisco-IOS-XE-interfaces-oper:interfaces/interface=GigabitEthernet0%2F1/statistics    | python3 -m json.tool
!
# That path, confirmed, is the one to subscribe to.

How much a path covers

A path naming a container returns everything beneath it, which on a device with many interfaces is a great deal per update. A path naming one leaf under one list entry returns one value.

The volume follows directly from how far up the tree the path sits, multiplied by the frequency. That is the calculation to do before configuring anything, and it is the one that prevents the collector-overwhelming outcome described later.

The operational half of the model

Subscriptions are for state rather than configuration: counters, statuses, tables that the device maintains. The configuration half is available too and changes rarely, which makes reporting it on change a reasonable thing to want — a subscription that fires when somebody changes something.

That is a genuinely useful application and it is different from the usual one. Most subscriptions are for counters and most counters change constantly.

! A subscription: what, how often, where to, in what encoding
R1(config)# telemetry ietf subscription 101
R1(config-mdt-subs)# stream yang-push
R1(config-mdt-subs)# filter xpath /interfaces-ios-xe-oper:interfaces/interface/statistics
R1(config-mdt-subs)# update-policy periodic 3000
R1(config-mdt-subs)# encoding encode-kvgpb
R1(config-mdt-subs)# source-address 10.255.0.1
R1(config-mdt-subs)# receiver ip address 10.200.0.80 57000 protocol grpc-tcp
!
R1# show telemetry ietf subscription 101 detail

The encoding

Several are available and they trade size against self-description. A compact binary encoding is smallest and requires the collector to hold the schema. A self-describing binary encoding carries the field names and is larger. A text encoding is largest and easiest to inspect while developing.

The self-describing binary form is the usual choice: considerably smaller than text, and the collector does not need a schema distributed to it. The compact form is worth it only at volumes where the difference matters.

Naming the source

Which address the device uses for the connection. The collector identifies the device by it, so a device whose telemetry arrives from whichever interface routing chose will be identified inconsistently — and, on a collector that maps addresses to device names, may appear as two devices.

Pinning it to a loopback is the same argument as for every other exchange with a central system, and the same one line.

Several subscriptions per device

Different data has different natural frequencies. Interface counters at a short interval; environmental readings at a long one; configuration on change. Those are three subscriptions with three intervals rather than one compromise.

There is no reason to use one, and using one means either sampling slow-moving data pointlessly often or sampling fast-moving data too rarely to be useful.

Choice Options Usual answer
Path Anywhere in the model As specific as the question allows
Encoding Compact, self-describing, text Self-describing binary
Source address Whatever routing chooses, or pinned Pinned to a loopback
Number of subscriptions One, or several Several, by natural frequency
Verify the path with a single read firstA path that returns what you want from one ordinary request is a path that will work in a subscription. Developing the path that way takes seconds per attempt; developing it by configuring a subscription and waiting to see whether anything arrives does not.
Sub claimThe volume follows directly from how far up the tree the path sits multiplied by the frequency, which makes the path choice a capacity decision rather than only a question of what data you want.

Dial-Out or Dial-In?

Which end connects?

Either. In one arrangement the subscription is configured on the device and the device opens a connection to the collector, retrying if it fails. In the other the device runs a server and a collector connects and subscribes over that session, with the subscription defined by the collector and not stored on the device. The first works where the collector cannot reach in and keeps the configuration visible on the device; the second keeps it in one place and leaves the device showing nothing about what is subscribed.

A Deeper Dive into the Two Directions

The device connecting outward

Configured on the device, in the running configuration, where it is backed up, appears in a configuration comparison and can be audited. The device opens the connection and retries on failure, which means a collector restart is recovered from automatically.

Its cost is that the configuration is on every device. Changing what is subscribed means a configuration change across the fleet, which is what configuration management exists for and is still a change.

The collector connecting inward

The device runs a server and the collector opens a session and states what it wants. The subscription exists only for the life of that session and is not part of the device's configuration.

That makes changing it a collector-side operation, which is considerably lighter. It also means the device's configuration contains nothing about what is being collected from it, so an audit of a device cannot answer the question.

! Dial-in: the device runs a server and stores no subscription
R1(config)# gnmi-yang
R1(config)# gnmi-yang server
R1(config)# gnmi-yang port 50052
!
! With transport security, which is the arrangement to prefer
R1(config)# gnmi-yang secure-server
R1(config)# gnmi-yang secure-trustpoint TP-GNMI
R1(config)# gnmi-yang secure-port 9339
!
R1# show gnmi-yang state detail

Subscribing from the collector side

A client opens a session and describes what it wants: the paths, the mode, and for a streaming mode the interval or the on-change instruction. The device begins sending and continues until the session ends.

The modes available are a single response, a response on request, and a continuous stream. The stream is the interesting one and it subdivides into sampling at an interval and reporting on change.

# Subscribe from the collector, no device configuration involved
$ gnmic -a 10.1.1.1:9339 -u admin -p pass --skip-verify     subscribe --mode stream --stream-mode sample     --sample-interval 30s     --path "/interfaces/interface[name=GigabitEthernet0/1]/state/counters"
!
# Ask the device what it supports, first
$ gnmic -a 10.1.1.1:9339 -u admin -p pass --skip-verify capabilities

Which to choose

Outward where the collector cannot reach the devices, which covers branches behind address translation and anything across an untrusted path. Outward also where the subscription should be visible in the device's configuration for audit reasons.

Inward where the collector can reach everything and the ability to change subscriptions without a configuration change is worth more than device-side visibility. In practice large deployments frequently use both, for different data.

Transport security

Both directions can run with or without it, and telemetry carries operational detail about the network. Where the path is not a controlled management network, unprotected transport is exposing that detail to whatever is on the path.

Enabling it requires a certificate on the device and the corresponding trust on the collector, which is real work and is the reason it is frequently skipped. On a management network it is a defensible omission; elsewhere it is not.

Pitfall: subscriptions exist and no device configuration mentions them Symptom: a device is generating a substantial volume of telemetry and its configuration contains nothing about any subscription. An audit of what is being collected from the device cannot answer the question from the device. Cause: the subscriptions were established by a collector connecting inward, so they exist only for the life of those sessions and are not part of the configuration. Confirm: the device shows active sessions and subscriptions in the subsystem's own state, which is separate from the running configuration. Fix: query the subsystem's state rather than the configuration when auditing, and record on the collector side what each device is subscribed to — because the device will not tell you.
Question Device dials out Collector dials in
Works when the collector cannot reach in Yes No
Visible in the device configuration Yes No
Changing what is collected A configuration change per device Collector side only
Survives a collector restart Device retries Collector reconnects
Audit from the device Read the configuration Read the subsystem state
Sub claimA subscription established by a collector connecting inward does not appear in the device's configuration, so an audit of what is collected from a device has to query the subsystem's state rather than reading the configuration.

How Often, and What Does the Number Mean?

What is the unit?

Hundredths of a second, in the device-configured form. A value of one thousand is ten seconds and a value of ten is a tenth of a second. Reading it as seconds and entering ten produces a subscription sending ten updates per second, which on a broad path across many interfaces is a large multiple of what was intended. Nothing about the configuration looks wrong and the first symptom appears at the collector.

A Deeper Dive into Frequency

The arithmetic to do first

How many objects the path covers, multiplied by how many values each carries, multiplied by the rate. A path covering interface statistics on a switch with forty-eight ports, at a value that turns out to be ten times a second, is a great deal of data per device and a fleet's worth at the collector.

That calculation takes a minute and it is the difference between a deployment that works and one that is rolled back after overwhelming the collector.

Pitfall: the collector is overwhelmed shortly after a subscription is deployed Symptom: a telemetry collector begins dropping data or failing shortly after subscriptions are configured across a set of devices. The devices report no problem and their subscriptions are valid and connected. Cause: the periodic interval is expressed in hundredths of a second. A value entered as though it were seconds produces a rate one hundred times higher than intended, which on a broad path is an enormous volume. Confirm: read the configured value and divide by one hundred; that is the interval in seconds. Fix: correct the value, and calculate the expected volume before deploying a subscription across a fleet rather than after.

Choosing an interval

Fast enough to answer the question and no faster. Interface counters for capacity work are useful at thirty seconds and not obviously more useful at one. Something being watched during an incident may justify a second or two, temporarily.

The temptation is to collect everything as often as possible because the mechanism makes it easy. The constraint is the collector and the storage behind it, and both are considerably more expensive than the device's effort to send.

Different data, different intervals

Counters change continuously and benefit from regular sampling. Environmental readings change slowly and a minute is generous. Configuration changes rarely and should be reported on change rather than sampled at all.

Three subscriptions with three settings is the right structure and costs nothing beyond the configuration. One subscription at a compromise interval collects the slow data pointlessly often and the fast data too rarely.

! Three subscriptions, three natural frequencies
!
R1(config)# telemetry ietf subscription 101
R1(config-mdt-subs)# filter xpath /interfaces-ios-xe-oper:interfaces/interface/statistics
R1(config-mdt-subs)# update-policy periodic 3000        ! 30 seconds
!
R1(config)# telemetry ietf subscription 102
R1(config-mdt-subs)# filter xpath /environment-ios-xe-oper:environment-sensors
R1(config-mdt-subs)# update-policy periodic 30000       ! 5 minutes
!
R1(config)# telemetry ietf subscription 103
R1(config-mdt-subs)# filter xpath /cdp-ios-xe-oper:cdp-neighbor-details
R1(config-mdt-subs)# update-policy on-change            ! only when it changes

The interval in the collector-driven form

Expressed as a duration with a unit, which removes the ambiguity entirely. That is a genuine advantage of that arrangement and worth knowing when comparing the two.

It does mean the two forms of the same deployment express the same thing differently, which is worth noting in documentation so that somebody reading both does not convert one into the other incorrectly.

What the device costs

Assembling and sending an update is work, and a device with several subscriptions at short intervals across broad paths does a measurable amount of it. That is usually small relative to the device's capacity and it is not zero.

The check is the processor utilisation before and after deploying subscriptions, on one device, before doing it everywhere. If it moved noticeably, the paths or the intervals are too aggressive.

! Measure the device cost on one, before the fleet
R1# show processes cpu sorted | head 10
R1# show telemetry ietf subscription all brief
R1# show telemetry ietf subscription 101 detail | include Update|Encoding|State
!
! And what is actually being sent
R1# show telemetry connection all

Timestamps

Each update carries the device's own timestamp, which is what makes subscription data more trustworthy than polled data. It also means the device's clock matters: an unsynchronised device produces data that cannot be correlated with anything.

An authenticated time source is a prerequisite rather than a refinement here, for the same reason it is for logging, and more visibly because the timestamps are in every single update. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.

Data Changes Interval
Interface counters Continuously Tens of seconds
Protocol neighbour state Rarely On change
Environmental readings Slowly Minutes
Configuration Rarely On change
Something under investigation — Seconds, temporarily
Sub claimThe device's effort to send is small and the collector's effort to receive and store is not, which makes the interval a decision about the collector's capacity rather than about the device's.

What Does Reporting Only on Change Require?

What has to be true?

Three things. The model must support change notification for the node being subscribed to, which not every node does — and a subscription requesting it for one that does not is either rejected or quietly degraded. The value must actually change from time to time, or nothing is ever sent and the subscription looks broken. And the collector must tolerate long silences without concluding the device has gone.

A Deeper Dive into Change Notification

Why it is attractive

For anything that changes rarely, sampling is waste. A neighbour table, a protocol adjacency state, a configuration — sampling those every thirty seconds sends the same values a thousand times a day to report a handful of changes.

Reporting only the changes sends those few updates and nothing else, which is a reduction of orders of magnitude for exactly the data where knowing about the change quickly matters most.

The support question

Whether a given node can be reported on change is a property of the model and its implementation. Some nodes support it, some do not, and the list differs by platform and by software version.

The practical check is to configure it and read the subscription's state. A subscription that reports itself as invalid has been rejected; one that reports valid and delivers periodic updates has been degraded to sampling, which is the confusing case because it appears to be working.

Pitfall: an on-change subscription is delivering regular periodic updates Symptom: a subscription configured to report only on change delivers updates at a steady interval regardless of whether anything changed. The subscription reports itself as valid and connected. Cause: the node being subscribed to does not support change notification on this platform or software version, and the subscription has been satisfied by periodic sampling instead of being rejected. Confirm: the update timestamps are evenly spaced while the values are identical between them. Fix: confirm change notification support for that specific path before relying on it, and where it is unsupported use a long sampling interval deliberately rather than an on-change subscription that is not one.

The silence problem

A subscription reporting only on change sends nothing while nothing changes, which for stable data is most of the time. A collector treating silence as a failure will report the device as dead within whatever timeout it applies.

The arrangements that address this are a periodic heartbeat alongside the change reporting, where supported, or a collector configured to understand that silence is expected. The first is better because it distinguishes a healthy quiet subscription from a broken one.

What it is genuinely good for

Adjacency states, interface states, neighbour tables, configuration. Anything where the event is what matters and the steady value is not interesting. For those, change reporting delivers a notification within a second of the event rather than within a sampling interval, which is a meaningful operational improvement.

For counters it is meaningless, because they change constantly and a change subscription would send continuously.

! On change for state, periodic for counters. Never the reverse.
!
R1(config)# telemetry ietf subscription 110
R1(config-mdt-subs)# filter xpath /ospf-oper:ospf-oper-data/ospf-state/ospf-instance/ospf-area/ospf-interface/ospf-neighbor
R1(config-mdt-subs)# update-policy on-change
R1(config-mdt-subs)# encoding encode-kvgpb
R1(config-mdt-subs)# receiver ip address 10.200.0.80 57000 protocol grpc-tcp
!
R1# show telemetry ietf subscription 110 detail | include State|Update Policy

The initial update

A change subscription typically sends the current state once when it starts, so the collector begins with a complete picture rather than with nothing until something changes. Without that, a collector starting during a quiet period has no data at all for an unknown time.

Whether it happens depends on the implementation, and it is worth confirming rather than assuming — the difference between a collector that knows the current state and one that does not is significant during an incident.

Combining both

Two subscriptions to the same data, one reporting changes immediately and one sampling at a long interval, gives fast notification and a periodic confirmation that the subscription is alive and the state is what the changes implied.

That is a small amount of extra data for a meaningful increase in confidence, and it resolves the silence problem without depending on heartbeat support.

! Both, to the same data: fast notification plus a slow confirmation
R1(config)# telemetry ietf subscription 110
R1(config-mdt-subs)# filter xpath /ospf-oper:ospf-oper-data/.../ospf-neighbor
R1(config-mdt-subs)# update-policy on-change
!
R1(config)# telemetry ietf subscription 111
R1(config-mdt-subs)# filter xpath /ospf-oper:ospf-oper-data/.../ospf-neighbor
R1(config-mdt-subs)# update-policy periodic 30000     ! every 5 minutes, as a heartbeat
Data On change Why
Adjacency state Yes The event is the whole point
Interface state Yes Same
Neighbour tables Yes Changes are rare and significant
Configuration Yes Every change matters
Counters No They change constantly
Environmental readings No They drift continuously
Pair an on-change subscription with a slow periodic oneThe change subscription gives the notification and the periodic one proves the subscription is alive and confirms the state. It is a small amount of extra data and it removes the ambiguity between a quiet subscription and a broken one.
Sub claimA subscription requesting change notification for an unsupported node may be satisfied by periodic sampling rather than rejected, which produces something that appears to be working and is not what was asked for.

Why Is the Collector Receiving Nothing?

What is the order?

Three states, read in order. The subscription's own state, which is invalid if the filter names something the device does not implement. The receiver's state, which sits at connecting if nothing is listening or the path is blocked. And, once connected, whether data is being sent — which if it is and the collector still shows nothing points at the encoding or the collector's ability to decode it. Each state is one line of output and the order rules out everything below it.

A Deeper Dive into the Silence

An invalid subscription

The filter names a path the device does not implement, or a model absent from this software version. The subscription exists in the configuration and reports itself as invalid, and nothing is ever sent.

The fix is the same as everywhere else in this material: ask the device what it implements rather than adjusting the path by guesswork, and verify the path with a single read before subscribing.

! Three states, in order
!
R1# show telemetry ietf subscription all brief
  ID     Type     State     Filter type
  101    Config   Valid     xpath
  102    Config   Invalid   xpath
!
R1# show telemetry ietf subscription 101 receiver
  Address        Port    Protocol   State
  10.200.0.80    57000   grpc-tcp   Connecting
!
R1# show telemetry connection all

A receiver that will not connect

The device is trying and nothing is accepting. Either the collector is not listening on that port, or something in the path is discarding the connection, or the address is wrong.

The distinguishing test is an ordinary connectivity check from the device's source address to the collector's port. It succeeds and the problem is the collector; it fails and the problem is the path.

Connected and nothing arriving

The device reports a connection and the collector reports no data. That is an encoding question: the collector is receiving bytes it cannot interpret, and depending on its implementation it either logs an error or silently discards them.

Matching the encoding to what the collector expects resolves it. The self-describing binary form is the most widely handled and is the safe choice when in doubt.

Pitfall: the device reports a healthy connection and the collector logs nothing at all Symptom: the subscription is valid, the receiver state shows connected, the device's counters show data being sent, and the collector has no record of receiving anything from that device. Cause: the encoding the device is using is not one the collector can decode. Depending on the collector, undecodable data is discarded silently rather than reported. Confirm: the device's sent counters increment while the collector's received counters for that device do not. Fix: set the encoding to the self-describing binary form, which is the most widely supported, and confirm the collector's configuration names the same one.

The source address again

A collector mapping addresses to device names will not recognise a device whose telemetry arrives from an unexpected address, and may discard it or file it under an unknown device.

That produces the specific symptom of data arriving and not appearing under the device it came from, which looks like a collector problem and is a source address problem.

Time

Every update carries the device's timestamp. A device whose clock is wrong produces data that lands in the wrong place on a timeline, or that a collector rejects as too far from the present.

Checking the device's synchronisation is part of commissioning telemetry rather than a separate concern, because the data is worth very little if it cannot be correlated with anything else.

! The commissioning checks, all of them
R1# show telemetry ietf subscription all brief
R1# show telemetry ietf subscription 101 detail | include State|Encoding|Update
R1# show telemetry connection all
R1# ping 10.200.0.80 source Loopback0
R1# show ntp status | include synchronized|stratum
R1# show processes cpu sorted | head 5

What to monitor about the telemetry

The count of valid subscriptions against the count configured, which detects one becoming invalid after an upgrade. The receiver connection state, which detects a collector that went away. And the collector's own received-per-device figures, which detect the device that stopped sending without anything on the device indicating it.

The third is the important one and it requires the collector's cooperation, which is the same conversation as for every other collection mechanism.

After an upgrade

Model paths can change between software versions, so a subscription valid before an upgrade can be invalid after it. Nothing announces this, and the symptom is a device that quietly stops reporting.

Checking subscription validity as part of post-upgrade verification is one command and it catches the whole class. Without it, the discovery is made when somebody notices a gap in a dashboard weeks later.

Blueprint framing

The CCIE Enterprise Infrastructure v1.1 blueprint includes model-driven telemetry within its automation domain. What is examined is generally the subscription model — what is subscribed to, which direction the connection goes, and the difference between periodic and change-based reporting — rather than collector-side tooling.

State Means Look at
Subscription invalid The path is not implemented The device's model list
Receiver connecting Nothing is accepting Collector, or the path
Connected, nothing decoded Encoding mismatch Both ends' encoding
Data under an unknown device Source address not pinned The source setting
Data at the wrong time Clock not synchronised The time source
Stopped after an upgrade Model path changed Subscription validity
Sub claimA model path that changes between software versions turns a valid subscription into an invalid one with no announcement, which makes checking subscription validity a required post-upgrade step rather than an optional one.

Conclusion

Subscribing replaces the polling loop with a push timed by the device's own clock, which removes the per-request overhead and produces data timed by the thing being measured rather than by the thing measuring it. The subscription names a path into the same models everything else uses, so the efficient way to develop one is to verify the path with a single ordinary read and then subscribe to what returned the value you wanted.

The interval in the device-configured form is expressed in hundredths of a second, and reading it as seconds produces a rate a hundred times higher than intended. Nothing on the device objects — the first symptom is the collector failing, some time after the subscription was deployed across a fleet. Calculate the volume from the breadth of the path multiplied by the rate before deploying, because the device's effort to send is small and the collector's effort to receive and store is not.

Use change reporting for state and periodic sampling for counters, never the reverse, and confirm that the specific path supports change notification — a subscription requesting it for a node that does not may be satisfied by periodic sampling instead of being rejected, which looks like it is working. And when nothing arrives, read three states in order: the subscription's validity, the receiver's connection state, and whether data is being sent. Each is one line and the order eliminates everything below it. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.

Reference Notes

  1. RFC 8639 specifies subscriptions to YANG notifications, including the distinction between a subscription established by configuration and one established dynamically over a session.
  2. RFC 8639 describes subscription state and the conditions under which a subscription may be rejected or terminated.
  3. RFC 8641 specifies subscription to datastore updates, defining both periodic subscriptions with a configured period and on-change subscriptions.
  4. RFC 8641 states that on-change support is not required for every node and that a publisher may decline an on-change subscription for nodes it cannot support.
  5. RFC 8641 describes the option of sending the current state when a subscription begins, so that a receiver has a complete picture before any change occurs.
  6. RFC 8641 defines a periodic subscription's period as a number of hundredths of a second.
  7. RFC 7950 specifies YANG, in which the paths subscribed to are defined, and distinguishes configuration from state data.
  8. RFC 8525 defines the YANG library, by which a device reports which modules and revisions it implements — the authority for whether a subscribed path exists.
  9. RFC 8040 specifies RESTCONF, which addresses the same model paths and is therefore the convenient way to verify a path before subscribing to it.
  10. RFC 5905 specifies NTP, whose synchronisation determines whether the timestamps carried in telemetry updates can be correlated with other data.
  11. Cisco documentation describes configured telemetry subscriptions, including the stream, filter, update policy, encoding and receiver parameters.
  12. The CCIE Enterprise Infrastructure v1.1 unified exam topics include model-driven telemetry within the automation domain.