Back to Journal Migration Strategy

Cutover Nights Without the 2am Panic

Dual-write data patterns, pre-flight rehearsal drills, and the hard rollback decision tree I print and tape to the wall during production cutovers.

Server racks glowing in a dark data centre during a night cutover window

A calm cutover is an engineering achievement built on dry runs and predetermined exit criteria.

We’ve all lived through the classic migration nightmare: the clock hits 02:30 AM on a Saturday, a script fails halfway through database index rebuilding, the lead DBA hasn't slept in twenty hours, and nobody knows whether to push forward or try to restore the backup.

It doesn’t have to be this way. Over forty-eight enterprise platform migrations, I have refined a cutover framework that turns migration night into a boring, routine execution where everyone goes to sleep on time.

1. Shadow Dual-Write Before Cutover Day

The cardinal sin of data migrations is attempting to move all data in a single maintenance window. If your dataset exceeds 500GB, offline dump-and-load is a recipe for failure.

Instead, establish asynchronous continuous replication two to three weeks prior to cutover. New transactions write to the primary legacy database and mirror instantly to the cloud replica via change data capture (CDC) such as Debezium or AWS DMS. By the time cutover night arrives, the cloud database is already 99.99% synced, with only seconds of replication lag to flush.

canary-cutover-route53.tf
resource "aws_route53_record" "api_canary" {
  zone_id = aws_route53_zone.primary.zone_id
  name    = "api.company.com"
  type    = "A"

  weighted_routing_policy {
    weight = 10 # Start with 10% canary traffic
  }

  set_identifier = "cloud-canary"
  alias {
    name                   = aws_lb.cloud_alb.dns_name
    zone_id                = aws_lb.cloud_alb.zone_id
    evaluate_target_health = true
  }
}

2. The Mandatory Rule of Three Rehearsals

Never attempt a production cutover without running at least three full rehearsal drills in staging:

  • Dry Run 1 (The Clock Test): Time every step down to the second. Identify which schema migrations take 45 minutes instead of the estimated 5.
  • Dry Run 2 (The Chaos Test): Intentionally pull the plug on a database replica mid-sync to verify your runbook covers failure states.
  • Dry Run 3 (The People Test): Run the cutover with the exact engineers who will be on-call, with the architect acting purely as an observer.
The "Point of No Return" Rule

Before cutover starts, write down a hard timestamp (e.g. 03:00 AM). If primary data reconciliation is not 100% verified by that exact minute, you abort and roll back immediately. No debate, no heroics.

3. The Rollback Decision Tree I Tape to the Wall

In high-stress moments, human reasoning degrades. A printed decision tree removes debate:

  1. Is replication divergence > 5 seconds at T-minus 15? → Abort. Remain on legacy estate.
  2. Does canary HTTP 5xx rate exceed 0.1% in first 10 minutes? → Shift Route 53 weight back to 100% legacy.
  3. Are customer debit events failing reconciliation in Redis? → Trigger automated reverse replication and revert DNS TTL.

Pre-Cutover 24-Hour Checklist

Ryan Cole portrait

Written by Ryan Cole

Cloud Architect specializing in zero-downtime database cutovers and mission-critical cloud migrations for high-growth tech companies.

Plan a Cutover

Get Monthly Cloud Architecture Notes

One monthly email detailing real migration debriefs, architecture trade-offs, and tested scripts. Never spam.