Curated summary
Failure is inevitable: Learning from a large outage, and building for reliability in depth at Datadog
Datadog’s March 2023 outage exposed a fundamental weakness in its reliability strategy: although 40–50% of production Kubernetes nodes remained operational, customers experienced the platform as entirely unavailable. The incident showed that preventing every failure is impossible and that systems must instead continue delivering useful, accurate service when components fail. Datadog consequently began redesigning products around graceful degradation, prioritizing data preservation, fresh information, and partial results.
Lessons from the March 2023 Incident
- An unsupervised global update triggered a restart interaction that disconnected roughly 50–60% of production Kubernetes nodes.
- The web interface recovered quickly, but logs, metrics, alerts, traces, and other core features became unavailable.
- Pages loaded without displaying customer data, creating a nearly complete outage from the user’s perspective.
Limits of Traditional Root-Cause Analysis
- Datadog identified the legacy global security-update mechanism as the immediate trigger and disabled it.
- Fixing that mechanism alone could not address the broader class of failures caused by certificates, configuration changes, overloads, date-handling bugs, or other unexpected events.
- The company concluded that resilience requires reducing the impact of failures, not merely preventing one specific failure mode.
Why Partial Infrastructure Became a Total User-Facing Failure
- Datadog’s systems historically favored complete correctness over partial visibility.
- For example, metric queries could wait until all relevant tags were processed to avoid showing misleading values or triggering false alerts.
- During a large outage, this behavior created a “square-wave” failure: missing some data caused the system to show no data.
- Ordered queues could stall fresh results behind stuck work, retries could overload already-strained services, and node-specific processing could make surviving capacity ineffective.
- The underlying design assumption was that systems should either function fully or stop, rather than degrade while continuing to provide value.
Prioritizing Graceful Degradation
Datadog shifted from relying primarily on redundancy and “never-fail” architectures to explicitly designing for inevitable failures.
- Customer data should never be lost, even if delivery is delayed.
- Fresh, real-time data should take priority over stale backlog processing.
- Systems should provide partial but accurate results whenever possible instead of returning nothing.
Persistent Storage at the Start of Processing Pipelines
- The outage caused a limited but non-zero amount of irreversible customer data loss.
- Some pipelines acknowledged data before writing it to replicated storage, leaving unreplicated data only in memory or on a local disk.
- When a node failed, that data disappeared and could not be recovered through agent retries.
- After the node loss, surviving intake nodes also struggled to write to downstream replicated stores.
- Their memory and local-disk buffers eventually filled, causing additional data loss as the outage continued.
- Datadog therefore identified persistent intake storage as a key requirement for preserving data during large-scale failures.
The broader recommendation is to design systems not only to prevent outages, but also to remain useful during them: preserve every accepted event, prioritize current information, and expose accurate partial results instead of failing completely.
Related reading
Continue with another curated summary.
A one-line Kubernetes fix that saved 600 hours a year
Read originalHow we tracked down a Go 1.24 memory regression across hundreds of pods
Read originalAchieving relentless Kafka reliability at scale with the Streaming Platform
Read originalUnraveling a Postgres segfault that uncovered an Arm64 JIT compiler bug
Read original