Every industry has its migration ghost story: the three-year lift-and-shift that doubled costs, the cutover weekend that became a cutover month, the system that came back online missing a decade of edge cases. These stories are common enough that some companies stay on aging infrastructure out of pure fear — which is its own slow-motion failure.
The truth is that migration horror follows a recognisable pattern, and so does migration success. The difference is almost never tooling. It is sequencing, honesty about the existing system, and refusing the big-bang cutover in favour of moves that can be individually verified and individually reversed.
Why migrations go wrong
The classic failure opens with underestimating the current system — undocumented dependencies, forgotten integrations, batch jobs nobody owns but everybody needs. The second act is the all-at-once cutover: months of parallel work converging on one weekend where everything must go right simultaneously. The third act writes itself.
The other quiet failure is migrating without modernising judgement: hauling every inefficiency into the cloud unchanged, then discovering the new bill prices those inefficiencies hourly. Lift-and-shift has its place — but as a chosen step, not a default.
The pattern that works
Inventory first: map what exists, what talks to what, and what can be retired — the cheapest workload to migrate is the one you delete. Then move in slices, lowest-risk first, each slice behind a validated rollback plan you have actually rehearsed. Run old and new in parallel where the stakes justify it, compare outputs, and cut over when the data says so — not when the calendar does.
Zero-downtime is not a slogan; it is a set of techniques — replication, dual-running, traffic shifting — chosen per workload. The best migrations are the ones users never notice happened.
The payoff for doing it right
A well-run migration ends with more than relocated servers: current infrastructure that engineers can hire for, costs that track usage, security patched by default, and a platform that stops vetoing the product roadmap. Companies routinely describe the aftermath as "we stopped fighting our infrastructure" — which is the actual point.
Risk registers, rehearsed rollbacks, verified slices. It is not glamorous. It is how you become a company with a migration success story instead of a ghost story.