Latest Cisco, PMP, AWS, CompTIA, Microsoft Materials on SALE Get Now Get Now

JSON, XML, YAML and Jinja: The File Is Valid and Describes the Wrong Thing

Four things are usually listed together here and only three of them are the same kind of thing. Two structured formats and one human-friendly one are three ways of writing a tree of named values; they carry the same information and convert between each other with only edge cases lost. The fourth is a template engine, which is a different category entirely, and treating it as a fourth format is the source of most of the confusion.

The practical consequences follow from that. A conversion between the three structured formats is nearly free and worth doing whenever one is easier to work with than another. Producing configuration from a template is not a conversion at all — it is generating text, with no knowledge of what valid configuration looks like, which means a template will emit a syntactically perfect file describing something that cannot exist and nothing will object until a device does.

This article covers what these formats are actually encoding, where the three structured ones disagree, why the human-friendly one silently changes values, what a template engine is and why it is not a format, and how to validate any of it before it reaches a device. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.

Blog ClaimThree of these four are the same thing written differently and the fourth is not a format at all — a template engine emits text, so it will produce a syntactically perfect file describing a configuration that cannot exist, and nothing between it and the device will object.
The three structured formats encode the same tree of named values and differ in a handful of properties, while the template engine is a different category entirely — it produces characters, not structure.

What Are These Formats Encoding?

What is the common structure?

A tree. Named values, values that are themselves collections of named values, and ordered lists of either. Every one of the three structured formats writes that same tree with different punctuation, which is why converting between them is mostly mechanical and why learning one properly means most of the others come with it. Knowing that the tree is the thing, and the syntax is decoration, is the single most useful idea here.

A Deeper Dive into the Structure

Why the tree matters more than the syntax

A model describes the tree: which names exist, what type each value has, which are required, where lists appear. Both structured formats used by network devices encode that same model, and a device will accept either where it supports both.

So the question "which format" is usually answered by which interface is being used rather than by preference. The interesting question is always about the tree, and that is described by the model rather than by the format.

Where each is actually used

The session-based configuration protocol uses the markup format, for historical reasons and because its namespace mechanism is genuinely useful when several models contribute to one document. The web-style interface uses the object format, because that is what web interfaces use. Human-authored files — inventories, variables, playbooks — use the indentation-based format, because it is the only one of the three anybody wants to type.

That division is convention rather than necessity and it is stable enough to rely on. It also explains why all three appear in one workflow: a human writes the third, tooling converts it to the first or second, and the device receives whichever its interface expects.

# The same tree, converted, in two commands
$ python3 -c "import sys,yaml,json; json.dump(yaml.safe_load(sys.stdin), sys.stdout, indent=2)" < vars.yml
!
# And back the other way
$ python3 -c "import sys,yaml,json; print(yaml.safe_dump(json.load(sys.stdin), default_flow_style=False))" < data.json

Lists, which are where the formats differ most

An ordered collection is explicit in two of the formats and implicit in the markup one, where a repeated element name is a list. That difference is the main reason conversion between the markup format and the others is not perfectly reversible: a list of one is indistinguishable from a single item without consulting the model.

In practice that means automatic conversion from the markup format is reliable only when the model is available to say which names are lists. Conversion in the other direction is unambiguous.

Types

Two of the formats carry types — a number is a number, a boolean is a boolean. The markup format does not: everything is text, and the type comes from the model. That is not a defect, it is the design, and it means a device parsing the markup format validates against the model rather than against the syntax.

The consequence for scripts is that a value read from the markup format arrives as text and may need converting, where the same value from the object format arrives already typed.

Namespaces

The markup format has a built-in mechanism for saying which model a name belongs to, which matters when a document combines several. The object format achieves the same thing by convention, prefixing the name with the module.

Omitting it is a common failure with both: a request whose top-level name carries no module qualification is ambiguous, and the device rejects it with a message about an unknown element. Adding the prefix is the fix and it is easy to overlook when copying an example that had it.

! The module qualification is not optional
!
! Correct - the module is named
{"ietf-interfaces:interface": {"name": "Loopback0"}}
!
! Rejected - the device does not know which model this belongs to
{"interface": {"name": "Loopback0"}}
!
! The markup equivalent, with the namespace declared
<interface xmlns="urn:ietf:params:xml:ns:yang:ietf-interfaces">
  <name>Loopback0</name>
</interface>
Format Used by Types Namespaces
Markup Session-based configuration None — from the model Built in
Object Web-style interfaces Native By name prefix
Indentation-based Human-authored files Native, and guessed None
Learn the tree, not the syntaxAll three formats encode the same structure and the model describes it. Time spent understanding what the model says a valid document looks like transfers to every format; time spent memorising the punctuation of one does not.
Sub claimConversion from the markup format is reliable only with the model available, because a repeated element name is a list and a list of one is indistinguishable from a single item without it.

Where Do the Three Disagree?

What are the real differences?

Comments, which only two support and one does not. Attributes, which only the markup format has and which have no equivalent in the others. Duplicate names, which behave differently everywhere. Ordering, which the markup format preserves and the object format does not guarantee. And whitespace, which matters in exactly one of them and where a tab character is an error rather than indentation.

A Deeper Dive into the Differences

Comments

The object format has none. Not discouraged — absent from the specification, so a parser is entitled to reject them. That makes it a poor choice for a file a human maintains, and it is the main reason the indentation-based format exists for that purpose.

The workaround people reach for — a key called something like a comment — is data rather than a comment, and it will be sent to whatever consumes the file. Occasionally that matters.

Attributes

The markup format allows a value to be attached to an element rather than contained by it. That has no counterpart in the other two, so a conversion has to invent one, typically by prefixing the name.

Whatever convention is chosen, a round trip through another format and back does not reproduce the original document byte for byte. For network models this rarely matters because they use elements rather than attributes, but any tool converting arbitrary markup documents will hit it.

Duplicate names

Repeating a name within the same object is invalid in one format, a list in another, and accepted-with-the-last-winning in several parsers of the third. That inconsistency is worth knowing because a duplicated key in a hand-edited file is a common mistake and its effect depends on which parser reads it.

A linter catches it in the formats where it is an error and warns in the others. Running one is the answer.

# A duplicated key: silently keeps the last one in many parsers
interfaces:
  - name: GigabitEthernet0/1
    description: uplink
    description: to the core          # the first is discarded
!
# Lint before anything consumes it
$ yamllint vars.yml
$ python3 -c "import yaml,sys; yaml.safe_load(open('vars.yml'))"

Ordering

The markup format preserves the order of elements and the model may depend on it. The object format's members are conventionally unordered, so a parser is entitled to present them in any order and a comparison between two documents cannot rely on it.

The practical effect appears in comparisons: two equivalent documents in the object format may differ textually while describing the same thing. Comparing parsed structures rather than text is the reliable approach.

Whitespace and tabs

Indentation carries meaning in exactly one of the three, and a tab character is not valid indentation in it. An editor configured to insert tabs will produce a file that fails to parse with a message about a tab, which is at least clear.

Less clear is inconsistent indentation depth, which can produce a document that parses into a structure other than the one intended — a key nested one level too deep is valid and wrong.

Pitfall: a file parses successfully into the wrong structure Symptom: a variables file is accepted without error and the tooling behaves as though a setting is absent. The setting is plainly present in the file and correctly spelled. Cause: its indentation places it one level deeper than intended, so it is a member of the previous item rather than a sibling. The document is valid and describes something other than what was meant. Confirm: parse the file and print the resulting structure; the misplaced key appears under the wrong parent. Fix: correct the indentation, and make printing the parsed structure the routine check after editing rather than relying on the absence of a parse error.

Numbers

Large integers can lose precision in parsers that represent all numbers as floating point, which affects identifiers and counters more than anything else in network data. Where a number is really an identifier rather than a quantity, carrying it as text avoids the question entirely.

This is uncommon in practice and worth knowing because when it happens the symptom — an identifier that is almost right — is baffling.

What to use where

Human-authored files in the indentation-based format, for the comments and the readability. Anything a program generates or consumes in the object format, which has the fewest surprises. And the markup format wherever the protocol requires it, converted from a structure rather than assembled by hand.

Assembling the markup format by hand is where namespace errors come from. Building the structure and letting a library serialise it removes that class entirely.

Difference Markup Object Indentation-based
Comments Yes No Yes
Attributes Yes No No
Duplicate names A list Invalid Often last wins
Order guaranteed Yes No No
Whitespace significant No No Yes, and no tabs
Sub claimA document that parses without error may still describe something other than what was intended, which is why printing the parsed structure is the routine check rather than the absence of an error.

Why Does It Change Your Values?

What is happening?

The indentation-based format guesses the type of an unquoted value from its appearance. That is convenient for numbers and booleans written as such, and it misfires on values that merely look like them. A country code becomes a boolean. A version string becomes a number and loses a digit. An identifier with leading zeros loses them. Each of these is the format doing exactly what it was designed to do, and each produces a value that is not what was written.

A Deeper Dive into Type Guessing

The words that become booleans

A set of words are interpreted as true or false, and it is larger than just those two. Depending on the parser's version it includes short forms and affirmatives, which means an unquoted two-letter country code, a column of yes and no answers, or a value that happens to be one of those words arrives as a boolean.

The effect downstream is a configuration line rendering the word "False" where a name was intended, or a comparison against a string never matching because the value is not a string.

Pitfall: a value in a variables file becomes a boolean Symptom: a generated configuration contains the word "False" or "True" where a short text value was expected, or a comparison against a value in a variables file never matches despite the value being visibly correct in the file. Cause: the value was written unquoted and consists of a word the parser interprets as a boolean. Common cases are two-letter codes and affirmative words. Confirm: parse the file and print the type of that value; it is a boolean rather than text. Fix: quote it. As a habit, quote every value that is not unambiguously a number, which removes this entire class at the cost of two characters.

Numbers that lose digits

A version written as a decimal is read as a number, and a trailing zero after the decimal point disappears because numerically it is not there. A software version becomes a different string when rendered back, and comparisons against it fail in ways that look like a mistake elsewhere.

Anything that is an identifier rather than a quantity should be quoted for this reason. Versions, serial numbers, identifiers, anything with a leading zero, and anything containing more than one separator.

Leading zeros

A zero-padded identifier read as a number loses its padding, and in some parsers a leading zero signals a different base entirely, which changes the value rather than merely its formatting.

That second case is the dangerous one: an identifier that looks like it was padded becomes a different number, and the resulting configuration is applied to the wrong thing. Quoting prevents it.

# Every one of these is wrong, and all of them parse
country: no              # becomes False
version: 1.10            # becomes 1.1
vlan: 0012               # becomes 12, or worse, base-8
enabled: on              # becomes True
time: 12:30              # may become a number of seconds
!
# All of these are right, and cost two characters each
country: "no"
version: "1.10"
vlan: "0012"
enabled: "on"
time: "12:30"

The rule that covers all of it

Quote everything that is not unambiguously a quantity. A count is a number; an identifier is text even when it is made of digits. An interface name, a version, a code, a VLAN identifier, a serial number — all text.

Applied consistently this removes the entire class of problem with no downside beyond two characters per value. It is the single most valuable habit in working with this format.

Checking rather than trusting

Parsing a file and printing the resulting types, once, after editing it, catches everything in this section. It takes one line and it is the only reliable check, because the file looks correct in every case described here.

# One line, after any edit. It catches every case above.
$ python3 -c "
import yaml, json, sys
d = yaml.safe_load(open('vars.yml'))
print(json.dumps(d, indent=2))
for k, v in d.items():
    print(f'{k:<20} {type(v).__name__:<8} {v!r}')
"

Anchors and references

The format allows a value to be defined once and referenced elsewhere, which removes repetition in a large inventory. It is genuinely useful and it makes a file harder to read, because a value's definition may be a long way from its use.

Worth using for genuinely repeated blocks and worth avoiding for the sake of cleverness. A reader who cannot find where a value came from is a reader who will duplicate it. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.

Written as Parsed as Should be
no boolean false "no"
1.10 number 1.1 "1.10"
0012 number, padding lost "0012"
on / off boolean quoted
1500 number — correct unquoted is fine
Sub claimEvery misparse in this section produces a file that looks completely correct to a reader, which is why printing the parsed types after editing is the only check that works.

What Is a Template Engine, and Why Is It Not a Format?

What does it do?

It takes text with placeholders and a set of values, and produces text. That is the entire operation. It does not know what a configuration is, it does not validate anything, and it has no concept of the structure it is producing. So it will happily emit a file that is syntactically fine and describes a configuration that cannot exist, and nothing between it and the device will notice — because to everything in between, it is just text.

A Deeper Dive into Templating

The three constructs

Substituting a value, repeating a block for each item in a collection, and including a block conditionally. Everything else is refinement. A template using those three covers essentially every configuration generation task anybody does.

Filters — transformations applied to a value as it is substituted — are the fourth thing worth knowing, particularly the one that supplies a default when a value is absent, because that one prevents the most common failure.

# Three constructs cover essentially everything
interface {{ intf.name }}
 description {{ intf.description | default('unused') }}
{% if intf.address is defined %}
 ip address {{ intf.address }} {{ intf.mask }}
{% else %}
 no ip address
 shutdown
{% endif %}
{% for vlan in intf.vlans | default([]) %}
 switchport trunk allowed vlan add {{ vlan }}
{% endfor %}
!

The missing variable

A placeholder whose value is absent renders as nothing by default. The line keeps its command and loses its argument, producing something that is still text and is no longer a valid command.

Worse, a line that becomes empty may leave a preceding line as the last one in a block, changing what a following line applies to. That is how a template produces configuration that is valid and applies settings to the wrong object.

Pitfall: a rendered configuration is missing an argument and nothing objected Symptom: a generated configuration file contains a command with its argument missing — a description line with no text, an address line with one value instead of two. The template is correct and the rendering produced no error. Cause: a variable was absent from the data and the engine rendered it as an empty string, which is its default behaviour. The output is text and remains valid text. Confirm: the variable is absent from the data for that device while present for others. Fix: configure the engine to raise an error on an undefined variable, or supply a default for every substitution — the first is better because it fails loudly.

Making it fail loudly

The engine can be configured to raise an error when a placeholder has no value, rather than rendering nothing. That converts a silent malformation into an exception at generation time, which is exactly where you want it.

It is one setting and it should be the default in any template pipeline that generates configuration. Supplying explicit defaults for genuinely optional values then becomes a deliberate statement rather than an accident.

# Fail at render time rather than at the device
from jinja2 import Environment, FileSystemLoader, StrictUndefined

env = Environment(
    loader=FileSystemLoader("templates"),
    undefined=StrictUndefined,      # missing variable = error, not blank
    trim_blocks=True,                # control blocks do not leave blank lines
    lstrip_blocks=True,
    keep_trailing_newline=True)

out = env.get_template("interface.j2").render(intf=data)

Whitespace

Control blocks occupy lines in the template and, without adjustment, leave blank lines in the output. For configuration that is usually harmless and occasionally not, because some parsers treat a blank line as ending a block.

Two settings remove it: one that strips the newline after a control block and one that strips the indentation before it. Setting both once produces output that looks like configuration rather than like a rendered template, which also makes reviewing it far easier.

Reviewing the output

The rendered result is text and can be read before it is applied. That review is the safety mechanism this whole approach has, and skipping it means the first reader of the generated configuration is the device.

Rendering to a file, comparing against the device's current configuration, and reading the difference is a step that takes a minute and catches everything in this section.

# Render to a file, diff against reality, read it
$ python3 render.py --device R1 > /tmp/R1.generated
$ ssh R1 "show running-config" > /tmp/R1.current
$ diff -u /tmp/R1.current /tmp/R1.generated | less
!
# Only then apply it

Where the template ends and validation begins

A template can produce well-formed structured data rather than raw configuration lines, and that structure can then be validated against a model before being sent. That is a meaningfully safer arrangement than generating command lines, because the validation step exists.

Generating command lines remains common because it is what the devices accept most readily. Where the interface takes structured data, generating the structure and validating it is the better path.

Produces Validated by First reader
Command lines Nothing The device
Command lines, reviewed A human You
Structured data A model The validator
Turn on strict undefined behaviourOne setting turns a silently blank substitution into an error at render time. It is the highest-value single change available in a template pipeline, because the failure it prevents produces valid text and an invalid configuration — which nothing else in the chain will catch.
Sub claimA template engine produces characters rather than structure, so nothing between it and the device has any basis for objecting — which makes reviewing the rendered output the only safety mechanism the approach has.

How Do You Validate Before It Reaches a Device?

What are the layers?

Four, and each catches a different class. Lint the source file, which catches syntax and duplicated keys. Parse it and print the structure, which catches values that parsed into the wrong type or the wrong place. Validate the structure against a model, which catches names and types that the device would reject. And review the rendered output, which catches everything the previous three cannot know about. Each takes seconds and they are cumulative rather than alternatives.

A Deeper Dive into Validation

Linting

A syntax check plus a set of style and consistency rules. It catches the tab character, the duplicated key, the inconsistent indentation, and the line that is technically valid and probably not what was meant.

It is a single command and it belongs in whatever runs before a change is accepted. Running it only when something has gone wrong is missing the point.

# Layer 1: lint the source
$ yamllint -d relaxed inventory/
$ python3 -m json.tool payload.json > /dev/null && echo "JSON ok"
$ xmllint --noout config.xml && echo "XML well-formed"

Printing the parsed structure

The check that catches everything in the type-guessing section and the misplaced-key case. Parse, print, read. It cannot be automated into a pass or fail because the question is whether the structure matches the intent, which only a person knows.

Doing it after every edit to a file that matters is a habit worth building. It takes five seconds and it is the only thing that catches a document that is valid and wrong.

Validating against a model

Where a model exists — and for device configuration one does — a document can be checked against it before being sent. That catches an unknown name, a value of the wrong type, a missing required element and a value outside its permitted range.

This is the layer that turns "the device rejected it" into "the validator named the element", which is a considerable difference when the device is in production and the change window is short.

# Layer 3: validate against the model, before sending
$ pyang --strict -f tree ietf-interfaces.yang | head -20
$ yanglint -t config -s ietf-interfaces.yang candidate.xml
!
# Or against a schema, for non-model data
$ python3 -c "
import json, jsonschema
jsonschema.validate(json.load(open('payload.json')), json.load(open('schema.json')))
print('valid')
"

Reviewing the rendered result

The last layer and the only one that can catch a configuration that is structurally perfect and operationally wrong — the right syntax applied to the wrong interface, a policy referencing something that does not exist, a change that is valid and not what was intended.

Comparing against the device's current configuration rather than reading the generated file in isolation is what makes this effective, because the difference is short and the file is long.

Where these belong

In whatever runs before a change is applied, automatically, so that none of them depends on anybody remembering. The first three are pass-or-fail and can gate a change; the fourth needs a person and should be a required step rather than an automated one.

A pipeline with the first three and a human reading the difference is a meaningful control. One with none of them is a template engine connected directly to production.

# All four, in order, as a pre-change routine
set -e
yamllint inventory/                                   # 1. lint
python3 show_parsed.py inventory/host_vars/R1.yml     # 2. structure
python3 render.py --device R1 > /tmp/R1.generated     #    render
yanglint -t config -s models/*.yang /tmp/R1.xml       # 3. model
diff -u /tmp/R1.current /tmp/R1.generated             # 4. review

Version control

The source files belong in version control for the ordinary reasons and for one specific to this: the difference between two versions of a variables file is far more readable than the difference between two rendered configurations, and it says what somebody meant to change rather than what changed.

That makes the review of an intended change a review of a few lines rather than of a generated file, which is where review actually works.

What not to bother with

Elaborate schema definitions for data nobody else consumes. Validation of files that are read by exactly one script which would fail obviously on bad data. Converting between formats for the sake of consistency when each is being used where it belongs.

The effort belongs where a failure is expensive, which is anything generating configuration that reaches a device.

Blueprint framing

The CCIE Enterprise Infrastructure v1.1 blueprint includes data encoding formats within its automation domain. What is examined is the ability to read and produce these formats correctly rather than any particular tooling, and the published topic list is the authority on the scope.

Layer Catches Automatable
Lint Syntax, duplicates, tabs Yes
Print the structure Wrong type, wrong place No — needs a reader
Validate against a model Unknown names, bad types, ranges Yes
Review the difference Valid and not what was meant No — needs a person
Sub claimTwo of the four validation layers cannot be automated because they ask whether the result matches the intent, which is a question no checker has the information to answer.

Conclusion

Three of these are the same tree written with different punctuation, and understanding the tree transfers to all of them while memorising any one syntax transfers to none. The markup format carries no types and gets them from the model, has namespaces built in, and represents a list as a repeated element — which is why converting out of it needs the model to disambiguate. The object format has native types, no comments, and no guaranteed ordering. The indentation-based one is the only one anybody wants to type.

That last one guesses types from appearance, and the guess is wrong for anything that merely looks like a number or a boolean. A country code becomes false, a version loses a digit, a padded identifier loses its padding. Every one of those produces a file that looks completely correct, so the only check that works is parsing it and printing the resulting types. Quoting everything that is not unambiguously a quantity removes the whole class for two characters per value.

The fourth is not a format. A template engine fills in blanks and emits characters, with no idea what configuration is, so a missing value renders as nothing and leaves a command without its argument — still valid text, no longer a valid command, and nothing downstream has any basis for objecting. Turn on strict handling of undefined values so that fails at render time, and read the difference between the generated configuration and the device's current one before applying it. That review is the only safety mechanism this approach has. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.

Reference Notes

  1. RFC 8259 specifies JSON, defining objects, arrays and the scalar types, and provides no syntax for comments.
  2. RFC 8259 states that the names within an object should be unique and describes the varying behaviour of implementations when they are not.
  3. RFC 8259 notes that implementations vary in the precision with which they represent numbers, so interoperable use of large integers requires care.
  4. RFC 7950 specifies YANG, in which the structure, types and constraints of configuration and state data are defined.
  5. RFC 7950 distinguishes leaf, container, list and leaf-list nodes, which is the basis of how the same data is represented as repeated elements or as arrays.
  6. RFC 7951 specifies the JSON encoding of YANG-modelled data, including the rule that a name is qualified by its module where the module differs from the parent's.
  7. RFC 6241 specifies NETCONF, whose messages are XML documents and which relies on XML namespaces to identify the models contributing to a document.
  8. RFC 8040 specifies RESTCONF, which supports both XML and JSON encodings of the same modelled data.
  9. The W3C Namespaces in XML specification defines the namespace mechanism by which element names are qualified, which has no direct equivalent in JSON.
  10. RFC 7950 defines constraints including ranges, patterns and mandatory nodes, which a validator can check against a document before it is sent to a device.
  11. The YAML specification defines implicit typing of unquoted scalars, in which a value's type is resolved from its form, with the resolution rules differing between specification versions and implementations.
  12. The CCIE Enterprise Infrastructure v1.1 unified exam topics include data encoding formats within the automation domain.