Migrating a multi-site network to a VXLAN fabric is straightforward on paper and messy in practice, mostly because every site you're migrating already has its own undocumented history of quick fixes.
The Problem With 'Just Follow the Runbook'
A VXLAN migration plan assumes a known starting state: consistent VLAN numbering, documented uplinks, predictable spanning-tree behavior. Real sites that have grown organically over years rarely match that assumption. The first week of any multi-site migration isn't configuration — it's discovery, and skipping it is the single most common cause of an unplanned outage mid-cutover.
Across a set of sites that had never had a unified wireless architecture, the actual VXLAN and Catalyst WLC rollout took less engineering time than the discovery phase that preceded it. Every site had at least one undocumented VLAN doing something load-bearing that nobody remembered configuring.
Sequencing the Cutover
- Build the new fabric in parallel, never in place — the old network stays untouched until the new one is validated
- Migrate one site fully before touching the second, even if the plan says they're identical
- Keep a rollback path live for at least one full business cycle after cutover, not just overnight
- Treat every site's plan as a draft until it survives contact with that site's actual wiring
When the Plan Changes Mid-Project
On a multi-year, multi-site rollout, the network plan will change more than once — a site gets new hardware requirements, a department needs a segment nobody scoped, or a discovered legacy dependency forces a redesign of one hop in the fabric. The teams that handle this well aren't the ones with the most detailed original plan; they're the ones who built in enough slack to absorb a redesign without blowing the whole schedule.
The plan is a starting position, not a contract. Sites don't care what the diagram says they should look like.
What Zero Downtime Actually Requires
Zero-downtime wireless rollouts, department by department, aren't achieved by working faster — they're achieved by never cutting over a segment until traffic has been mirrored and validated on the new path first. It's slower per-site and faster overall, because you're not spending the next three days firefighting a cutover that should have been caught in validation.