Migrating Data Ingestion Systems at Meta Scale (opens in new tab)
Meta rebuilt its hyperscale MySQL data ingestion system to improve reliability, efficiency, and data-langing latency. The migration moved workloads from customer-owned pipelines to a simpler, self-managed warehouse service and ultimately transitioned 100% of jobs. Success depended on staged validation, continuous data comparison, and fast rollback mechanisms.
Why Meta Migrated
- The system incrementally moved several petabytes of social graph data from MySQL into Meta’s data warehouse each day.
- This data supports analytics, reporting, machine learning, and product development.
- The legacy architecture became increasingly unstable as data-landing requirements grew stricter.
- Customer-owned pipelines worked at smaller scales but became difficult to manage reliably at hyperscale.
Migration Success Criteria
Each job had to meet defined requirements before advancing:
- Data correctness: Old and new systems had matching row counts and checksums.
- Landing latency: The new system performed at least as well as the legacy system.
- Resource usage: Compute and storage consumption did not regress.
- Critical-table requirements: Additional criteria were agreed upon with dependent teams.
Three-Phase Migration Lifecycle
Shadow Phase
- New-system shadow jobs ran against the same production sources as existing jobs.
- Their output was written to separate shadow tables.
- Row counts and checksums were continuously compared with production data.
- Compute and storage requirements were measured before production rollout.
- Once validated in pre-production, shadow jobs were tested in production.
Reverse Shadow Phase
- The new system began writing to the production table.
- The legacy system continued running, but wrote to a shadow table.
- This preserved continuous comparison between both systems.
- If discrepancies appeared, Meta could quickly restore the old system without rebuilding its configuration.
Migration Cleanup
- Both systems continued to be monitored for mismatches.
- After validation, the legacy shadow job was removed.
- The new system became the sole production pipeline.
Data Quality and Debugging Tooling
- Meta built tooling to compare corresponding table partitions from the two systems.
- Comparisons included row counts, checksums, and example rows responsible for mismatches.
- Mismatch records and debugging details were logged to Scuba for real-time analysis.
- Hourly queries helped engineers identify root causes and determine whether issues were already known.
- The same tooling remains part of post-migration release validation.
Rollout and Rollback Controls
- Both systems used change data capture (CDC), with internal full-dump and delta tables feeding customer-facing target tables.
- Because CDC builds new data from previously landed data, an existing defect could propagate after migration.
- Meta therefore emphasized:
- Detecting problems before they reached data consumers.
- Stopping further propagation quickly during rollback.
- The reverse-shadow design provided early quality signals and preserved a ready-to-use legacy pipeline for rapid recovery.
Meta’s migration demonstrates that large-scale infrastructure changes are safest when treated as controlled, observable lifecycle transitions rather than one-time cutovers. Parallel execution, automated data validation, explicit resource checks, and reversible rollouts enabled the company to migrate the entire workload while protecting downstream consumers.