Comparing wide area options usually turns into a comparison of transports — a carrier service against the internet, one circuit type against another, cost per megabit. Those matter and they are not the decision. Every option on the list is a different answer to one question: who decides which path a packet takes.
Buy a carrier's routed service and the carrier decides, within your routing policy but out of your sight. Build an overlay over the internet and your routing protocol decides, with everything visible and everything your responsibility. Deploy a controller-driven design and a policy you wrote once decides, per application, on measurements taken continuously. Those are three different operating models, and the transport underneath is almost incidental to how the network will feel to run.
This article covers what the options actually are, who owns the path decision in each, what each one costs when it fails, how to combine them without producing something worse than either, and which of these decisions are hard to reverse. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.
What Are You Choosing Between?
What are the options?
Four, and most real networks use two of them. A carrier's routed service, where the carrier participates in your routing and provides connectivity between sites as a product. An overlay you build across whatever transport is available, usually the internet. A controller-driven overlay, which is the same idea with the path decision moved into a central policy. And local breakout at each site, which is not an alternative to the others but a decision about what does not need to cross the wide area at all.
A Deeper Dive into the Options
The carrier's routed service
You hand the carrier your routes and it delivers them to your other sites. Any site reaches any other without you configuring anything per pair, the capacity is contracted with a guarantee attached, and the path between sites is the carrier's business rather than yours.
What you give up is visibility and agility. A performance problem inside the carrier's network is something you can observe the effects of and not the cause, and a change to how traffic is treated is a conversation rather than a configuration. What you gain is somebody contractually obliged to care.
The overlay you build
Tunnels between your own devices over whatever transport exists, with your routing protocol deciding paths and your encryption protecting the traffic. The transport becomes a commodity: anything that carries packets between two of your routers will do.
Everything is yours, which is the advantage and the cost. You can see every path and measure it, and nobody is obliged to fix the underlying transport when it degrades. The operational model is the familiar one — routers, routing protocols, configuration per device.
The controller-driven overlay
The same tunnels, with the path decision lifted out of the routing protocol into a policy evaluated against continuous measurements of each path. Traffic for a given application takes the path that currently meets the requirement stated for it.
That is a genuinely different capability and it introduces a genuinely new dependency. The controllers are part of the network now: forwarding continues without them, and changes do not.
Local breakout
Traffic destined for the internet leaves at the site rather than crossing the wide area to a central exit. That removes a large and growing share of the traffic from the wide area entirely, which changes the sizing of everything else.
The cost is a security boundary at every site instead of one in the middle, which is a substantial operational and licensing question and the reason the decision is not obvious despite the bandwidth argument being overwhelming.
WHAT THE DECISION ACTUALLY DECIDES
Not: which transport is cheaper per megabit
But: who chooses the path, and what happens when it degrades
carrier service the carrier chooses. You observe effects.
your overlay your routing chooses. You see everything.
controller-driven a policy chooses, per application, on
measurements. You wrote the policy once.
local breakout decides what never crosses the wide area
Why most networks use two
A private service for what needs a guarantee and an internet-based overlay for everything else, or two internet transports with the overlay across both. Single-transport designs are unusual because the failure of that transport is the failure of the site.
Which two, and how traffic is divided between them, is the substance of the design and is covered later. The pairing itself is nearly universal.
THE FOUR OPTIONS, AS A SHOPPING LIST
carrier routed service
buy: connectivity between sites, with a guarantee
own: your edge routers and what you advertise
accept: no visibility past the handover
your own overlay
buy: transport, from anyone
own: tunnels, routing, encryption, measurement
accept: nobody is obliged to fix the transport
controller-driven overlay
buy: a platform, and transport from anyone
own: a policy, and the sites
accept: a dependency that freezes change when it is gone
local breakout
buy: a security boundary at every site
own: that boundary, per site
accept: it is a security decision, not a network one
What the requirements should decide
Whether a contractual guarantee is needed for anything, which points at a carrier service for that traffic. Whether per-application path choice is required, which points at a controller. Whether sites must reach each other directly, which every option supports differently. And what the operations team can run, which is the constraint that decides more of these than anybody admits.
| Option | Strongest when | Weakest when |
|---|---|---|
| Carrier routed service | A guarantee is genuinely required | Agility matters, or cost does |
| Your own overlay | You want visibility and control | Nobody is obliged to fix the transport |
| Controller-driven | Per-application path choice | A new platform dependency is unwelcome |
| Local breakout | Most traffic is internet-bound | A security boundary per site is costly |
Who Owns the Path Decision?
Why does that matter most?
Because it decides what you can do when something degrades, how a change is made, and what skills the team needs. A carrier owning it means a performance problem is a ticket. Your routing owning it means a performance problem is yours to diagnose and yours to route around. A policy owning it means the network already routed around it and you find out afterwards. Those are three different jobs for the same team.
A Deeper Dive into Ownership
When the carrier owns it
Your routers hand traffic over and it emerges at the far site. Between those two points the path, the treatment and the recovery are the carrier's. You influence it only through what you advertise and what the contract says.
The operational consequence is a clean division: problems are either inside your network or inside theirs, and establishing which is the first diagnostic step. That is simpler than owning everything, and it depends entirely on the carrier being responsive.
When your routing owns it
Paths are chosen by metrics you set, failover is your routing protocol's convergence, and quality is whatever the underlying transport delivers. Everything is visible and everything is your problem.
The strength is that a decision can be changed in an afternoon. The weakness is that the routing protocol chooses on topology rather than on quality: a path that is up and performing badly is still the best path, and nothing in the protocol notices.
WHAT A ROUTING PROTOCOL CANNOT SEE
It chooses on metric, which is a property of the topology.
It does not choose on:
current delay current loss current jitter
whether this path is good for voice right now
Adding probes and tracking bridges part of the gap - one
object per path, driving the metric. That is the manual
version of what a controller does continuously.
When a policy owns it
Each path is measured continuously, each application is given a requirement, and traffic takes whichever path currently satisfies it. That closes the gap the routing protocol has and it does so per application rather than per destination.
The change is operational as much as technical. A path decision is no longer a metric on a device; it is a statement in a policy applying everywhere, which is a different way of working and a different set of things to get wrong.
The manual version
Probes measuring each path, tracked objects driving route preference, and a policy directing specific traffic accordingly. That reproduces a useful fraction of the controller-driven behaviour using mechanisms that already exist.
It is genuinely workable for a handful of sites and it does not scale, because each site's probes, objects and policies are configured individually. The point at which that stops being manageable is roughly the point at which a controller starts being worth its dependency.
THE MANUAL VERSION, PER SITE
ip sla 10 / icmp-echo <far side> source-interface <primary>
ip sla 20 / icmp-echo <far side> source-interface <backup>
track 10 ip sla 10 reachability delay down 3 up 60
track 20 ip sla 20 reachability delay down 3 up 60
route-map VOICE-OUT permit 10
match ip address VOICE
set ip next-hop verify-availability <primary> 10 track 10
set ip next-hop verify-availability <backup> 20 track 20
Workable at ten sites. Unmanageable at two hundred.
That crossover is where a controller earns its dependency.
What the team has to be able to do
Under a carrier service: run a routing relationship with a provider and manage a contract. Under your own overlay: run tunnels, routing and encryption across an unreliable transport. Under a controller: run a platform, understand a policy language, and diagnose a network where the forwarding decision is not in the device you are logged into.
The third is the largest change and it is frequently underestimated, because it is presented as a simplification. It is a simplification of the configuration and a complication of the diagnosis.
Who owns it after a failure
The question worth asking during design: when the wide area is degraded at three in the morning, who is expected to act and with what information. A carrier service has an answer involving somebody else's support organisation. An overlay has an answer involving your team and complete visibility. A controller has an answer involving your team, complete visibility and a policy that may already have handled it.
Whichever is chosen, the answer should be written down, because it determines what monitoring and what training are needed.
| Owner | A degraded path is | You need |
|---|---|---|
| The carrier | A ticket | Evidence, and a contract |
| Your routing | Yours to detect and route around | Probes, tracking, and time |
| A policy | Already handled, probably | The skill to read why |
What Does Each Cost When It Fails?
Which failures matter?
Three shapes. A transport that is down, which every design handles by having a second one. A transport that is up and degraded, which is the failure that distinguishes the options — plain routing does not notice it, probes and tracking notice it slowly, a controller notices it immediately. And a failure of whatever makes the design work: a carrier's support organisation, your own overlay's control plane, or the controllers.
A Deeper Dive into Failure
The transport that is down
The easy case. Every design covers it with a second transport and the difference is only in how quickly. Routing convergence, probe-driven tracking or controller-measured steering all move traffic; they differ by seconds rather than by capability.
This is the failure everybody designs for and it is the one least likely to be the source of complaints, because it is unambiguous and it is handled.
The transport that is degraded
The failure that decides between the options. A circuit carrying traffic with elevated loss and delay is, to a routing protocol, a working path. Traffic continues to use it, applications suffer, and nothing in the network reports a fault.
Plain routing has no answer. Probes with tracking have one and it is coarse — the path is usable or it is not, for all traffic. A controller has a per-application answer and moves only the traffic that cares.
When the carrier's support is the failure
A degraded carrier service is only as recoverable as the carrier's responsiveness, and that is the risk the contract is supposed to address. It frequently does not, because the remedy is a credit rather than a repair.
The design answer is a second transport not from that carrier, which restores the ability to route around a problem you cannot fix. That is also why single-carrier designs with two circuits from the same provider are weaker than they appear.
When the controllers are unavailable
Forwarding continues — the edge devices have their policy and their measurements and keep applying them. What stops is change: new sites cannot be brought up, policy cannot be altered, and the visibility the design depends on is gone.
That is a manageable failure mode and it must be understood and stated. A team expecting a controller failure to be an outage will over-engineer around it; a team expecting nothing to happen will under-plan for the change freeze it actually causes.
When your own overlay's control plane fails
An overlay depends on its own resolution and routing. A hub failure, a resolution failure, or a routing problem in the overlay produces effects your team owns completely, with complete visibility and no external dependency.
That is the trade in its clearest form: nothing to escalate, and nobody else to rely on.
THE THREE FAILURES, AGAINST THE THREE DESIGNS
down degraded control plane
carrier service second ticket, wait their support
transport organisation
your overlay second probes, or yours - visible,
transport not at all and yours alone
controller-driven second automatic, forwarding
transport per app continues,
changes do not
The middle column is the one that decides the design.
What to measure regardless
Loss, delay and variation across every transport, continuously, independently of whatever the design does with the numbers. That is worth having under a carrier service as evidence for a ticket, under an overlay as the input to tracking, and under a controller as a check on what the controller believes.
It is also the only way to answer the question the business asks after a bad week, which is whether the transport was actually worse or whether something changed on your side.
Stating the failure behaviour in the design
For each failure, what happens, how long it takes, who acts and with what information. Four columns, one row per failure, and it takes an hour.
Its value is that it exposes the failures nobody had considered — almost always the degraded transport and the control plane — while the design can still accommodate them. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.
| Failure | Carrier service | Your overlay | Controller-driven |
|---|---|---|---|
| Transport down | Second transport | Second transport | Second transport |
| Transport degraded | A ticket | Probes, or nothing | Automatic, per application |
| Control plane lost | Their problem | Yours, visible | Forwarding continues, change freezes |
| Everything measured? | Do it anyway | Required | Do it anyway |
How Do You Combine Them?
What is the usual arrangement?
Two transports with one overlay across both, so that the overlay is transport-independent and either can carry everything. The alternative — a private service carrying some traffic natively and an overlay carrying the rest — produces two networks with two routing designs, two failure behaviours and two sets of things to understand. The first is simpler and is nearly always the better answer.
A Deeper Dive into Hybrid Designs
One overlay, two transports
Tunnels over both, the same routing inside them, and the transport chosen per path rather than per application boundary. The result is one network whose paths happen to traverse different underlying services.
That keeps the routing design singular, which is the property worth protecting. It also means a site losing one transport loses capacity rather than losing a category of connectivity.
ONE OVERLAY, TWO TRANSPORTS
Branch
Tunnel0 over the carrier service preferred
Tunnel1 over the internet backup, or active
One routing design inside both. One address plan.
One failure story: a transport loss is capacity loss.
Against:
Branch
routed natively into the carrier service design A
overlay to the hub over the internet design B
Two routing designs, two policies, two failure stories,
and traffic whose behaviour depends which one it landed in.
The two-network arrangement
Some traffic natively over a private service and the rest over an overlay. It arises naturally because the private service already existed and the overlay was added, and it is rarely chosen deliberately.
The cost is that every question has two answers. How does traffic fail over? Depends which network. What is the quality? Depends. Where is the policy applied? Two places. That accumulates into an operational burden nobody budgeted for.
Active-active or active-backup
Using both transports simultaneously doubles the available capacity and means both are continuously exercised, so a failure of one is not the first time it has carried anything. Using one as a standby keeps the traffic path predictable and leaves half the capacity idle.
The argument for active-active is strong and its cost is that traffic distribution becomes a thing to reason about — asymmetry, per-application placement, and what happens when one side's capacity is halved.
Where local breakout fits
It is orthogonal and it changes the sizing of everything else. A site sending most of its traffic to services on the internet, backhauling all of it to a central exit, is buying wide area capacity to carry traffic that never needed to be there.
The decision is not really about the network. It is about whether a security boundary can be operated at every site, which is a licensing, staffing and policy question. The network consequence — a large reduction in wide area traffic — is the easy part.
WHAT LOCAL BREAKOUT CHANGES
Before: branch -> wide area -> central exit -> internet
wide area sized for internet traffic plus internal
After: branch -> internet directly
wide area sized for internal traffic only
The bandwidth argument is usually overwhelming.
The decision is not about bandwidth - it is about whether
a security boundary can be operated at every site.
Keeping the routing design singular
Whatever the transports, one routing design, one addressing plan, one policy point. That is the discipline that keeps a hybrid manageable, and every departure from it should be justified specifically.
The common departure is a site brought up differently because of a constraint at that site, which then becomes a permanent exception nobody documented. One site configured unlike the others is the thing that makes a fleet-wide change impossible later.
The migration question
Most hybrids are arrived at rather than designed, by adding a transport to an existing service. That is fine and it is worth taking the opportunity to redesign the whole thing rather than bolting the new transport onto the old design.
Adding an overlay across both transports, and then removing the native routing over the private service, produces the singular design. Leaving both produces the two-network arrangement by default.
| Arrangement | Routing designs | Failure stories | Verdict |
|---|---|---|---|
| One overlay, two transports | One | One | The answer |
| Native plus overlay | Two | Two | Arrived at, not chosen |
| Active-active | One | One, with distribution | Worth the reasoning |
| Active-backup | One | One | Simpler, half idle |
| Local breakout | Orthogonal | Adds a boundary per site | Decided by security, not network |
Which Decisions Are Hard to Reverse?
What is the ordering?
Two of these decisions are measured in years and the rest in weeks. The carrier and the contract, because reversing it means new circuits to every site on somebody else's delivery schedule. And the controller platform, because reversing it means replacing the edge device at every site. Everything else — which overlay protocol, which routing design, active-active or not, even local breakout — is a configuration project. The analysis should be weighted accordingly.
A Deeper Dive into Reversibility
The carrier commitment
A multi-year contract with circuits delivered to every site. Changing carrier means ordering, installing and migrating every one of them, on a schedule that is not yours, with a period during which both are running and both are paid for.
That makes carrier selection the decision to spend the most analysis on, and it argues for transport arrangements that reduce the dependency — an overlay that treats the transport as a commodity means the carrier can be changed without redesigning anything above it.
The platform commitment
A controller-driven design ties the edge devices, the controllers and the policy language together. Leaving it means replacing the devices at every site, rebuilding the policy in whatever replaces it, and migrating sites one at a time.
That is a similar magnitude to a carrier change and it is frequently entered into with considerably less analysis, because it is presented as a technology choice rather than as a commitment. The question to ask explicitly is what leaving would involve, before joining.
WEIGHT THE ANALYSIS BY WHAT REVERSING COSTS
Years to reverse:
the carrier and the contract
the controller platform
Months:
local breakout (a security boundary per site)
Weeks:
which overlay protocol
active-active or active-backup
the routing design inside the overlay
Spend the analysis on the first group. The rest can be
decided quickly and changed later if they were wrong.
What reduces the dependency
An overlay that treats the transport as a commodity. If the design works over any transport that carries packets between your routers, then the carrier is a supplier rather than an architectural component, and changing one is a procurement exercise rather than a redesign.
That is the strongest long-term argument for building the overlay yourself or choosing a controller design that is genuinely transport-independent. It converts a years-long decision into a months-long one.
The decisions that feel permanent and are not
The overlay protocol, the routing design inside it, whether traffic is active-active. Each feels significant and each is a configuration change across the estate, which is weeks of work and no procurement.
Recognising this prevents the common pattern of spending months deliberating a reversible choice while the irreversible one — a three-year contract — is settled by whoever handled the procurement.
What to write down
For each decision: what it commits to, what reversing it would involve, and what evidence would change the answer. That last one is the useful part, because a decision recorded with the evidence that would overturn it can be revisited rationally.
Without it, a decision becomes a fact about the network that nobody remembers making and nobody feels able to question.
ONE ROW PER DECISION, IN THE DESIGN DOCUMENT
Decision: carrier routed service for site interconnect
Commits to: 3-year contract, circuits at 47 sites
Reversing: reorder and migrate all 47, 9-12 months,
overlap costs for the duration
Would change sustained failure to meet the agreed targets,
if: or internet-based paths measuring comparably
for two consecutive quarters
The last row is what makes it revisitable rather than
a fact nobody remembers choosing.
THE QUESTION TO ASK BEFORE JOINING, NOT AFTER
"If we decide in three years that this was wrong,
what exactly do we do?"
carrier service reorder circuits at every site,
9-12 months, overlap costs throughout
controller platform replace the edge device at every site,
rebuild the policy, migrate site by site
your own overlay reconfigure the edges. Weeks.
A design document with that answer in it for each option
is worth more than any amount of feature comparison.
The sunk cost problem
A long commitment creates pressure to keep using it beyond the point where it makes sense, because the cost is already committed. That is worth naming in the design document at the time, so that a later review is about what is best now rather than about justifying what was spent.
Stating the review point explicitly — a date, and what will be assessed — is the practical form.
Blueprint framing
The CCIE Enterprise Infrastructure v1.1 blueprint covers wide area design within its design domain. What is examined is generally the reasoning — what each option provides, what it costs, and which requirements point at which — rather than a specific product, and the published topic list is the authority on scope.
| Decision | Reversing costs | Analysis it deserves |
|---|---|---|
| Carrier and contract | Years | Most of it |
| Controller platform | Years | Most of it |
| Local breakout | Months, and a boundary per site | Substantial, mostly security |
| Overlay protocol | Weeks | Modest |
| Active-active or backup | Weeks | Modest |
| Routing inside the overlay | Weeks | Modest |
Conclusion
The options differ in who decides the path, and everything operational follows from that. A carrier service means the decision is theirs and a degraded path is a ticket. Your own overlay means the decision is your routing protocol's, which chooses on topology and cannot see quality at all. A controller-driven design moves the decision into a policy evaluated against continuous measurement, which closes that gap per application and introduces a platform the network now depends on.
Every design handles a transport that is down, and they differ by seconds. The failure that actually discriminates is the transport that is up and degraded — loss and delay on a circuit a routing protocol considers perfectly good. That is the failure users report, it is the one plain routing has no answer to, and choosing between the options is largely choosing how you intend to handle it.
Combine them with one overlay across both transports rather than one network natively and another over an overlay, because a singular routing design is what keeps a hybrid manageable and the alternative gives every operational question a "depends which". And weight the analysis by what reversing each decision costs: the carrier contract and the controller platform are measured in years, everything else in weeks, and the common pattern is to deliberate the reversible choices at length while the irreversible ones are settled by whoever handled the procurement. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.
External Links
- RFC 4364 — BGP/MPLS IP Virtual Private Networks
- RFC 2332 — NBMA Next Hop Resolution Protocol
- RFC 4301 — Security Architecture for the Internet Protocol
- RFC 4116 — Accountability and Multihoming
- RFC 2681 — A Round-trip Delay Metric for IPPM
- RFC 3393 — IP Packet Delay Variation Metric
- Cisco Learning Network — CCIE Enterprise Infrastructure
Reference Notes
- RFC 4364 specifies BGP/MPLS IP VPNs, in which a provider participates in the customer's routing and delivers any-to-any connectivity between customer sites as a service.
- RFC 4364 describes the provider edge to customer edge relationship, through which a customer's influence over the provider's path selection is limited to what it advertises.
- RFC 2332 defines NHRP, the resolution mechanism that allows an overlay to establish direct paths between sites without per-pair configuration.
- RFC 4301 defines the IPsec security architecture, which provides the confidentiality an overlay applies to traffic crossing an untrusted transport.
- RFC 4116 discusses multihoming, including the dependence of traffic distribution on routing policy when more than one path exists.
- RFC 2681 defines a round-trip delay metric, one of the measurements used to assess whether a path is meeting an application's requirements.
- RFC 3393 defines IP packet delay variation, the measurement most relevant to real-time traffic and one a routing protocol does not consider.
- RFC 3393 notes that delay variation is a property of the current path conditions rather than of the topology, which is why topology-based path selection cannot account for it.
- Cisco design guidance for wide area architectures describes transport-independent overlays and the separation of the overlay routing from the underlying transport.
- Cisco design guidance describes application-aware routing, in which path selection is driven by measured path characteristics against per-application requirements.
- Cisco design guidance describes direct internet access at branch sites and the security considerations that accompany moving the boundary to each site.
- The CCIE Enterprise Infrastructure v1.1 unified exam topics include wide area design within the design domain.