Migrating a campus to a fabric architecture is presented as a network project and it is not one. The fabric itself is well documented, largely automated, and the part that goes according to plan. What consumes the schedule is everything around it: the period during which both networks exist, the component joining them that appears in nobody's diagram, and an identity system that has to be working before the fabric delivers anything the old network did not.
That last point is the one worth stating first. The reason to build a fabric is usually segmentation — the ability to say which groups may reach which, independently of where anybody is plugged in. That depends entirely on knowing who each device belongs to, which depends on an identity deployment that is frequently further from ready than anybody has assessed. Without it the result is one large segment with a new control plane, which is most of the cost and none of the benefit.
This article covers what is actually being migrated, what has to be true before starting, how the two worlds coexist while it happens, the order things should move in, and why migrations stall. It is written for the lab rather than for the written exam, and sits alongside the rest of the CCIE Enterprise Infrastructure lab certification track.
What Are You Actually Migrating?
What changes?
Three things, and only one of them is the network. The forwarding changes, from routing and switching directly to an overlay with a control plane that maps hosts to locations. The policy model changes, from addresses and VLANs to groups that travel with the device. And the operating model changes, from configuring devices to expressing intent in a controller. The third is the largest change and the one least often planned for, because it is not a technical migration at all.
A Deeper Dive into the Change
The forwarding change
Traffic is encapsulated between edge devices and the control plane answers the question of where a given host currently is. That decouples addressing from location, which is the property everything else depends on: a device keeps its address wherever it connects.
Mechanically this is well trodden and largely automated. It is also the part everybody focuses on, which is why migrations are planned as network projects and then stall on something else.
The policy change
Instead of expressing access rules in terms of addresses and the VLANs those addresses live in, devices are assigned to groups and rules are written between groups. The rule follows the device regardless of where it connects and regardless of what address it has.
This is the actual benefit and it is entirely dependent on something knowing which group each device belongs to. That something is the identity system, and its readiness is the single largest determinant of whether the project delivers.
The operating model change
A controller expresses intent and configures the devices. A change made directly on a device is no longer a fix, it is drift, and the controller will either overwrite it or flag it — and either way the team's habitual way of working no longer applies.
That is a genuine change to how people do their jobs and it needs the same treatment as any such change: training, a period of adjustment, and an explicit decision about what remains permissible to do directly. Skipping that produces a controller-managed network that people keep configuring by hand, which is the worst of both.
THREE MIGRATIONS IN ONE PROJECT
forwarding overlay, with a control plane that locates hosts
well documented, largely automated, goes to plan
policy groups instead of addresses and VLANs
depends entirely on identity being ready
operating the controller is the source of truth
a change to how people work, rarely scheduled
Plans that address only the first one stall on the other two.
What does not change
The addressing, if the migration is done properly. Applications. What users experience, other than during the change itself. Those are worth stating explicitly because a migration that changes them has taken on a much larger project.
In particular, preserving addressing is a decision worth insisting on, because renumbering during a migration turns a switch replacement into a change affecting every device on it.
Whether to do it at all
The honest question, asked early. If the driver is segmentation, it is worth establishing what segmentation is actually required and whether a simpler approach reaches it. If the driver is a hardware refresh that happens to make it possible, the fabric is a choice rather than a consequence.
That question is uncomfortable partway through a project and cheap at the start. A design document stating what the fabric is for, in terms of a requirement, is the thing that makes the later decisions tractable.
The size of what is being taken on
A new control plane, a new policy model, a new operating model, a controller platform, an identity deployment, and a coexistence period covering all of it. Each is manageable and together they are a programme rather than a project.
Estimating it as a network migration is the most common planning error, and the schedule that results is wrong by a factor that only becomes visible once the identity work starts.
| Migrating | Difficulty | Usually planned for |
|---|---|---|
| Forwarding | Moderate, well documented | Yes — this is the plan |
| Policy model | High — depends on identity | Partially |
| Operating model | High — it is a people change | Rarely |
| Addressing | Should not change at all | Sometimes changed unnecessarily |
What Has to Be True First?
What are the prerequisites?
Five, and they are usually less ready than assumed. An identity system with real policy in it, not just deployed. Hardware and software that support the roles being asked of them. An underlay whose maximum frame size accommodates the encapsulation, which is almost never already done. An addressing plan that can survive the migration without renumbering. And a device inventory that is accurate, which is the one everybody believes they have.
A Deeper Dive into the Prerequisites
Identity, which is the long pole
The fabric assigns devices to groups, and something must decide which group. That decision comes from the identity system, which means it needs to be deployed, integrated with the directory, populated with policy that reflects how the organisation actually works, and trusted enough to enforce.
Each of those is substantial and the last is the hardest, because it requires knowing what every device on the network is — including the ones nobody remembers. Organisations that have not previously run identity-based access control routinely underestimate this by a large margin.
The frame size
Encapsulation adds bytes to every packet, so the underlay must carry frames larger than the endpoints send. If it does not, large packets fail while small ones pass, which produces the familiar selective failure: connections establish, transfers stall, and a ping succeeds.
Raising it across the underlay is straightforward and it must be consistent on every link. It is also almost never already done, and it is a prerequisite rather than a tuning step — a fabric on an underlay with the default size will appear to work during light testing and fail under real traffic.
! The underlay must carry more than the endpoints send
UNDERLAY(config)# interface range TenGigabitEthernet1/0/1 - 4
UNDERLAY(config-if-range)# mtu 9100
!
! Consistent on every underlay link. Verify rather than assume.
UNDERLAY# show interfaces TenGigabitEthernet1/0/1 | include MTU
UNDERLAY# ping 10.0.0.2 size 9000 df-bit
!
! Small packets passing proves nothing here.
Hardware and software
Each role has requirements, and a campus assembled over a decade will have devices that cannot fill some of them. Establishing which devices can be which role, before planning the phasing, prevents a plan that assumes a switch can do something it cannot.
Where a refresh is needed, it belongs in the programme's scope and budget rather than being discovered during it. That discovery is a common cause of a stalled phase.
The addressing
The migration should not renumber anything, which requires the existing addressing to be usable as it is. Where a subnet spans multiple switches, or where addressing is scattered in a way the fabric's model does not accommodate, some rationalisation is needed and it is better done before than during.
Assessing this early is cheap. Discovering it during a floor migration is a stopped change window.
The inventory
Every device on the network, what it is, how it is addressed, and whether it can authenticate. The gap between what the inventory says and what is actually connected is where the surprises are, and it is largest in exactly the places migration happens — printers, building systems, laboratory equipment, things installed by other departments.
Discovering that gap during a floor migration is expensive. Discovering it during a survey beforehand is a list of items to plan for.
THE SURVEY THAT SAVES THE SCHEDULE
For every access port with something on it:
what is it, and who owns it
does it authenticate, or can it
is it statically addressed
does it need to be in the same subnet as something else
can it tolerate a brief outage
The rows that are blank are the rows that stop a change window.
Filling them in beforehand is the cheapest work in the programme.
Verifying the underlay carries the frames
Before anything is migrated, prove a full-size packet crosses every underlay path without fragmenting. The verification is a large packet with fragmentation disallowed, sent between the devices that will carry encapsulated traffic.
A failure here is silent later: the fabric comes up, control plane registrations succeed, small flows work, and larger transfers stall in a way that looks like an application problem.
! Prove it before the first user moves, not after
EDGE-1# ping 10.0.0.9 size 9000 df-bit repeat 5
Type escape sequence to abort.
Sending 5, 9000-byte ICMP Echos to 10.0.0.9, timeout is 2 seconds:
Packet sent with the DF bit set
!!!!!
Success rate is 100 percent (5/5), round-trip min/avg/max = 2/3/5 ms
EDGE-1# show interfaces TenGigabitEthernet1/0/1 | include MTU
MTU 9100 bytes, BW 10000000 Kbit/sec, DLY 10 usec,
! A row of dots on ONE path is the whole finding.
A readiness gate
Stating the prerequisites as a gate, with a measure for each, prevents the pattern of starting the network work while the identity work is still in design. The most useful measure is the proportion of endpoints the identity system can classify confidently.
Below a threshold, the migration delivers a control plane and no policy. Above it, the policy model starts to be real. That single number is the most useful thing to track in the whole programme.
| Prerequisite | Measure | If unmet |
|---|---|---|
| Identity with policy | Proportion classified confidently | One large segment |
| Hardware and software | Devices meeting each role | A refresh mid-project |
| Underlay frame size | A large packet crossing | Selective failure under load |
| Addressing | Subnets usable unchanged | Renumbering during migration |
| Inventory | Ports with a known owner | Surprises, per floor |
How Do the Two Worlds Coexist?
What joins them?
A border, which is where the fabric meets everything else, and a device beyond it that leaks between the fabric's virtual networks and the shared services both worlds use. The second is the one that appears in nobody's first diagram and turns out to be central: every virtual network is separate by design, so anything shared — identity, name resolution, addressing, monitoring, printing — has to be deliberately reachable from each of them.
A Deeper Dive into Coexistence
The border
The point where fabric traffic leaves and external traffic enters. It carries the fabric's virtual networks outward as separate routing contexts, which preserves the separation across the boundary.
That is straightforward and it produces the requirement that follows: if the virtual networks remain separate outside the fabric, something has to join them where they need joining.
The device that leaks
Separate routing contexts that need to reach common services require a device participating in each and leaking routes between them selectively. That device is doing the work that makes the segmentation usable, and its configuration is where the policy about what may reach what actually lives during migration.
It is frequently discovered late, sized wrong, and placed somewhere inconvenient. Planning it as a first-class component — with capacity, redundancy and a policy model — is the single most useful correction to a typical first design.
! The device joining the virtual networks to shared services
FUSION(config)# vrf definition CORP
FUSION(config-vrf)# address-family ipv4
FUSION(config)# vrf definition GUEST
FUSION(config-vrf)# address-family ipv4
!
! Each virtual network arrives as its own context
FUSION(config)# interface GigabitEthernet0/1.101
FUSION(config-subif)# encapsulation dot1Q 101
FUSION(config-subif)# vrf forwarding CORP
FUSION(config-subif)# ip address 10.200.1.2 255.255.255.252
!
! And shared services are leaked in, deliberately and selectively
FUSION(config)# router bgp 65000
FUSION(config-router)# address-family ipv4 vrf CORP
FUSION(config-router-af)# import map SHARED-SERVICES-ONLY
Keeping a subnet in both worlds
The mechanism that makes a migration tolerable. A subnet can exist in the fabric and in the legacy network simultaneously, so a switch can be moved into the fabric while the devices on it keep their addresses and keep talking to devices still outside.
Without it, moving a switch means renumbering everything on it, which converts a switch replacement into a change affecting every user on that switch and requiring their applications to cope. With it, a floor migration is a switch swap.
Verifying a migrated host still reaches the old world
After the first switch moves, the check that matters is not that the host is in the fabric — the controller will say so — but that it still reaches a device that has not moved, and a shared service across the boundary, with its original address.
Doing that from the host, to a specific unmigrated peer and a specific shared service, is worth more than any controller status page, because it exercises the path through the border and the leaking device rather than reporting on the fabric alone.
! Confirm the host is where the control plane thinks it is
EDGE-1# show device-tracking database interface Gi1/0/12
Network Layer Address Link Layer Address Interface State Time
ARP 10.20.14.37 0050.56a1.3c9d Gi1/0/12 REACHABLE 31 s
! And that the address did NOT change across the migration
! 10.20.14.0/24 exists in both worlds during the overlap.
! From the host: a peer that has NOT moved, same subnet
C:\> ping 10.20.14.52
Reply from 10.20.14.52: bytes=32 time=2ms TTL=128
! From the host: a shared service, across the border
C:\> ping 10.99.1.10
Reply from 10.99.1.10: bytes=32 time=4ms TTL=124
! The second one is the test of the leaking device.
How long coexistence lasts
Longer than planned, always, because the last few devices are the difficult ones. A plan assuming a short overlap will find itself maintaining two networks for a year, and the design should be built for that rather than treating it as a failure.
Designing the coexistence properly — documented, monitored, with clear ownership — is better than designing a fast migration and improvising the long tail.
What has to work across the boundary
Identity, because devices in both worlds authenticate against the same system. Name resolution and addressing services. Monitoring. Any application whose clients are migrating in stages while its servers stay put.
Each of those is a dependency crossing the boundary and each is a thing to verify explicitly before the first migration rather than discover during it.
THE CROSS-BOUNDARY DEPENDENCY LIST
identity devices in both worlds, same system
address assignment relays reaching the same servers
name resolution same servers, from both sides
monitoring both worlds visible to the same tooling
applications clients migrating, servers staying
Verify each one from inside the fabric BEFORE the first
user floor moves. Each is a change window if discovered later.
Wireless during the transition
Wireless can traverse the fabric without being integrated into it, which allows the wired migration to proceed without the wireless design changing. Integrating it properly is a second phase.
Doing both at once is possible and it combines two sets of unknowns in the same change window. Separating them costs a phase and makes each one diagnosable. Reading this once is not the same as being able to do it under time pressure, which is what repetition against realistic CCIE lab practice scenarios is for.
| Component | Role during coexistence | Commonly |
|---|---|---|
| Border | Carries the virtual networks outward | Planned |
| The leaking device | Joins them to shared services | Discovered late |
| Subnet in both worlds | Avoids renumbering | Sometimes omitted |
| Wireless | Over the top, then integrated | Attempted at once |
| Cross-boundary services | Must work from both sides | Verified during, not before |
In What Order?
What is the sequence?
Build the fabric alongside and migrate nothing. Move one floor of ordinary devices and leave it for a week. Move the rest floor by floor. Bring wireless over the top, then integrate it. Turn policy on in observation mode and enforce one group at a time. And leave until last everything that is statically addressed, unusual, or not understood — because that population is where the schedule goes.
A Deeper Dive into Sequencing
Build first, migrate nothing
The fabric, the border, the leaking device and the controller, all working, with nothing on them. That proves the infrastructure independently of any user impact, and it is the phase where problems are cheap.
Resisting the temptation to migrate something during this phase is worth doing, because a problem with a user on it is a different kind of problem.
One floor, then wait
A single floor of ordinary user devices, migrated and then left alone for a week. The waiting is the point: problems that appear immediately are found in the change window, and the interesting ones appear over days as different applications are used.
Compressing this is the most common schedule decision and it moves the discovery of those problems to a point where several floors have the same fault.
THE PILOT, AND WHAT THE WEEK IS FOR
Day 0 migrate one floor. Everything works. Declare nothing.
Day 1-2 the applications people use daily
Day 3-4 the weekly ones - reporting, backups, batch jobs
Day 5-7 the ones nobody mentioned until they failed
The faults that matter are in the second and third rows,
and compressing the wait moves them to a point where
five floors already have them.
Ordinary devices first
Laptops and desktops obtaining addresses dynamically and authenticating normally. They are the largest population, the best understood, and the most tolerant of a brief interruption.
Migrating them first also builds the operational familiarity that everything afterwards depends on, on the population where a mistake is most recoverable.
Policy last, and gradually
Classification running and reporting what it would do, without enforcing. That produces the evidence about what is actually on the network and how it would be classified, which is the input to deciding whether enforcement is safe.
Then one group at a time, starting with a group whose access requirements are narrow and well understood. Enforcing everything at once on a population that has never been classified is how a migration becomes an outage.
POLICY, IN THE ONLY ORDER THAT WORKS
1. Classify, enforce nothing. Read the reports for a month.
2. Find the devices classified as unknown. Identify every one.
3. Enforce for ONE group, chosen because its access is narrow
and well understood. Leave it a fortnight.
4. Next group. Repeat.
Enforcing everything at once, on a population never
classified before, is how a migration becomes an outage.
What goes last
Statically addressed devices. Devices that cannot authenticate. Anything belonging to another department. Anything nobody can identify. Applications with unusual network requirements.
That population is small and consumes a disproportionate share of the schedule, because each item is its own investigation. Identifying it during the survey and planning it as its own phase is better than meeting it floor by floor.
Rollback
Per switch, back to the legacy configuration, provided the legacy path was left available. That requires not decommissioning the old distribution connectivity until the migration is genuinely complete, which costs some ports and buys the ability to undo a floor.
A migration with no rollback is one where every change window must succeed, which is a considerably more stressful project and produces worse decisions at two in the morning.
Declaring it finished
When the last device is in the fabric, the coexistence mechanisms are removed, the legacy path is decommissioned and the policy is enforced for every group. Each of those is a step and the project is not finished until all four are done.
Projects frequently stop after the migration of the easy population and leave the rest permanently, which produces exactly the two-network arrangement the migration was supposed to end.
| Phase | Population | Risk |
|---|---|---|
| Build, migrate nothing | None | Lowest — do it properly |
| One floor, wait a week | Ordinary devices | Contained |
| Floor by floor | Ordinary devices | Manageable |
| Wireless | All | Separate phase |
| Policy, group by group | All | Highest if rushed |
| The difficult population | Static, unusual, unknown | Consumes the schedule |
Why Do Migrations Stall?
What are the causes?
Four, in order of frequency. Identity not being ready, which leaves the policy model unusable and the benefit undelivered. The difficult population, which is small and consumes the schedule. The operating model change, which was not planned for and produces a controller-managed network people configure by hand. And a design covering the destination but not the coexistence, which means the long overlap is improvised.
A Deeper Dive into Stalling
Identity, again
The first and largest cause. A programme that scheduled the network work and the identity work in parallel reaches the point of enforcing policy with classification coverage too low to enforce anything, and stops there — with the fabric built and the benefit unavailable.
The correction is to gate the network work on a classification measure. That feels like delaying the visible part of the programme and it is the difference between delivering segmentation and delivering a control plane.
The difficult population
The devices that do not authenticate, are statically addressed, belong to somebody else, or cannot be identified. Each requires an investigation, a conversation with an owner, and frequently a decision about whether to accommodate it or replace it.
Scheduling this as a phase with its own duration, rather than expecting it to absorb into the floor-by-floor work, is what keeps the programme's end date meaningful.
The operating model
People continuing to configure devices directly because that is how they work, producing drift the controller flags or overwrites. The result is a team fighting the tooling and a network whose state nobody is confident about.
The correction is training and an explicit decision about what may still be done directly, made before the migration rather than after the first drift incident. It is a people problem and it responds to being treated as one.
Coexistence that was not designed
A design describing the destination, with the transition left as an implementation detail. The overlap then lasts a year and is run on improvisation, with the leaking device sized by guesswork and the cross-boundary dependencies discovered one at a time.
Designing the coexistence explicitly — what exists in both worlds, for how long, who owns it, how it is monitored — is the correction and it is a week of work at the start.
What to track
Classification coverage, which gates everything. Ports migrated against ports total. The size of the difficult population and its trend. And the number of groups under enforcement, which is the actual measure of delivered benefit.
The last is the one to report upward, because ports migrated measures activity and groups enforced measures outcome. A programme reporting only the first can be ninety percent complete and have delivered nothing.
THE FOUR NUMBERS WORTH REPORTING
classification coverage gates everything else
ports migrated measures activity
difficult population measures remaining risk
groups under enforcement measures delivered benefit
A programme reporting only the second can be ninety percent
complete and have delivered nothing anybody asked for.
What a healthy programme looks like
Identity ahead of the network work rather than beside it. A survey completed before the first floor. A coexistence design with an owner. A pilot floor left alone for a week. Policy in observation for a month before anything is enforced. And a difficult-population phase with its own schedule.
None of that is about the fabric, which is the point: the fabric is the documented part and everything listed here is the part specific to the organisation.
When to stop
If classification coverage is not improving and the segmentation requirement cannot be met, the honest options are to fix that first or to stop. Continuing produces a completed migration that delivered a new control plane, which is a expensive way to arrive where you started.
Saying so is uncomfortable and it is cheaper than the alternative. A programme with a stated purpose can be assessed against it; one without becomes an activity that finishes when the ports run out.
Blueprint framing
The CCIE Enterprise Infrastructure v1.1 blueprint covers software-defined access within its design and architecture domains. What is examined is generally the reasoning — the roles, how the fabric joins the rest of the network, and what coexistence requires — rather than a migration methodology, and the published topic list is the authority on scope.
| Cause | Symptom | Correction |
|---|---|---|
| Identity not ready | One large segment | Gate on classification coverage |
| The difficult population | Schedule consumed | Its own phase, its own duration |
| Operating model unplanned | Drift, and a team fighting tooling | Training and an explicit decision |
| Coexistence not designed | A year of improvisation | A week of design at the start |
| Reporting activity not outcome | Complete, and nothing delivered | Report groups enforced |
Conclusion
Three things migrate and only one of them is the network. The forwarding moves to an overlay, which is documented and largely automated and goes to plan. The policy model moves from addresses to groups, which is the actual benefit and depends entirely on an identity system that can classify devices with confidence. And the operating model moves to a controller, which is a change to how people work and is the one nobody schedules.
The migration is a coexistence project. Both networks exist for longer than planned, the subnet has to live in both so that a floor migration is a switch swap rather than a renumbering, and the device that leaks between the virtual networks and the shared services is central to the whole arrangement while appearing in nobody's first diagram. Designing that period explicitly — what exists in both worlds, for how long, owned by whom — is a week of work that replaces a year of improvisation.
And the thing to measure is not ports. A programme can migrate every port on schedule and deliver a new control plane with every device in one group, which is most of the cost and none of the reason it was funded. Gate the network work on classification coverage, report groups under enforcement rather than ports moved, and be willing to say early that the identity work has to come first — which is uncomfortable at the start of a programme and considerably cheaper than discovering it at the end. More CCIE Enterprise Infrastructure material — labs, protocol breakdowns and study guides — is collected on the SPOTO CCIE site.
External Links
- RFC 6830 — The Locator/ID Separation Protocol (LISP)
- RFC 7348 — Virtual eXtensible Local Area Network (VXLAN)
- RFC 4364 — BGP/MPLS IP Virtual Private Networks
- RFC 4459 — MTU and Fragmentation Issues with In-the-Network Tunneling
- IEEE 802.1X — Port-Based Network Access Control
- RFC 2072 — Router Renumbering Guide
- Cisco Learning Network — CCIE Enterprise Infrastructure
Reference Notes
- RFC 6830 specifies LISP, in which an identifier is separated from a locator so that a host's address no longer determines its position in the topology.
- RFC 6830 describes the mapping system that resolves an identifier to its current locator, which is the control plane function a fabric depends on.
- RFC 7348 specifies VXLAN, the encapsulation used to carry Layer 2 frames across a Layer 3 underlay, and notes the overhead it adds to each frame.
- RFC 7348 states that the underlay must accommodate the additional header, which is why the underlay maximum frame size must be raised before deployment.
- RFC 4459 describes the MTU and fragmentation problems created by in-network tunnelling, including the selective failure of large packets when the path cannot carry them.
- RFC 4364 describes the separation of routing contexts and the controlled leaking of routes between them, which is the mechanism by which separate virtual networks reach shared services.
- RFC 4364 describes route targets as the means of controlling which contexts import which routes, which is how selective leaking is expressed.
- IEEE 802.1X defines port-based network access control, the mechanism by which a device is authenticated and, from that authentication, assigned to a group.
- IEEE 802.1X describes the dependence of the authorisation decision on an authentication server, which is why the identity deployment gates the policy model.
- RFC 2072 describes the difficulty of renumbering hosts, which is why preserving addressing across a migration is worth the mechanism required to achieve it.
- Cisco software-defined access design guidance describes the border and the external device required to join virtual networks to shared services, and the handoff options available during migration.
- The CCIE Enterprise Infrastructure v1.1 unified exam topics include software-defined access within the design and architecture domains.