
01The 2am problem
A migration plan can look complete at kickoff: waves sequenced, dependencies mapped and the maintenance window booked. The harder question appears at 2am when replication is behind and the window closes at 6: who decides whether to continue, pause or roll back, and against which evidence?
Technical and operational decisions must therefore be designed together. The authorised decision owner must be reachable, the abort criteria must be visible to the bridge and the team must know how long the tested fallback takes. Without those elements, schedule pressure can quietly become the decision process.
A rollback that has never been rehearsed is not a plan. It is a hope with a diagram.
02Write the abort criteria before the plan
Write the rollback section before the forward path. It forces three commitments: a named decision owner who is reachable throughout the window; hard abort criteria expressed in measurable conditions—replication lag beyond a threshold, verification failures above a count or elapsed time past a gate; and a fallback rehearsed against representative data with its duration recorded.
That last point deserves emphasis. A rollback that has never been rehearsed is not a plan—it is a hope with a diagram. Rehearsal replaces an assumption with evidence: the sequence, access, dependencies and measured recovery time the decision owner needs when the window is under pressure.
03The go/no-go gate is not a formality
Every wave should pass a deliberate gate before traffic moves: parity verified, rollback confirmed ready, decision owner present. Teams skip this when the evening is going well, and that is precisely when it matters — the gate is cheap when things are fine and priceless when they are not.
If your migration partner cannot show you their abort criteria and their rollback rehearsal results for the previous wave, you do not have a migration plan. You have a schedule.