Curated summary
Scaling down to speed up: How we improved efficiency of live process metrics by 100x
Datadog redesigned its real-time Processes and Containers pipeline to avoid collecting high-frequency metrics that users never see. By limiting 2-second collection to hosts actively viewed and using standard 10-second data for sorting, the company reduced real-time traffic by over 100x, cut infrastructure costs by 98%, and lowered Agent resource usage. The approach also improved scalability without sacrificing the live investigation experience.
Original Real-Time Collection Model
- Datadog Agents normally collect process and container metrics every 10 seconds.
- When a user opened a live Processes or Containers view, all hosts in that tenant switched to 2-second collection.
- This supported near-real-time monitoring similar to
htop, but across distributed infrastructure. - As tenants grew, the pipeline had to process millions of processes per second, even though users typically viewed only around 50 processes or containers.
- Live sorting required keeping all tenant data in memory on a single server, limiting horizontal scaling and forcing vertical scaling.
Refocusing on User-Visible Data
- Most collected metrics were never displayed to users.
- Datadog determined that real-time collection only needed to be enabled for hosts running the processes or containers currently in view—up to roughly 50 hosts per user.
- Internal telemetry suggested this could reduce traffic by more than 100x.
- This required tracking active host subscriptions and updating them as users navigated the product.
- Because sorting occurred every 10 seconds, it did not need 2-second data. Datadog switched live views to use the existing 10-second metrics, aligning live and historical sorting logic.
Host Subscription Filtering
- A proof of concept added host subscriptions to the live data servers.
- Servers filtered Kafka payloads and discarded data for hosts without active subscriptions.
- This immediately reduced:
- Memory usage by 85%
- CPU usage by 33%
- The improvement came from storing fewer live metrics and processing fewer incoming payloads.
- The prototype confirmed that filtering preserved product behavior while simplifying sorting.
Moving Filtering Earlier in the Pipeline
- Late filtering improved live data servers but still left unnecessary work for the rest of the system and customer-side Datadog Agents.
- Datadog therefore planned to propagate subscription state to the intake service.
- Live data servers publish users’ active host sets over Kafka once per second.
- The intake service consumes this information and decides which hosts should activate 2-second process and container metric collection.
- This allows real-time collection to be restricted to hosts users are actively investigating while maintaining responsive live views.
Datadog’s redesign demonstrates that real-time systems scale more effectively when they prioritize data users can actually see. Filtering at intake, limiting high-frequency collection to subscribed hosts, and reusing standard-resolution data for sorting provide a simpler and more economical architecture without eliminating live functionality.
Related reading
Continue with another curated summary.
How we built reliable log delivery to thousands of unpredictable endpoints
Read originalHow we scaled fast, reliable configuration distribution to thousands of workload containers
Read originalTimeseries indexing at scale
Read originalIntroducing Husky, Datadog's third-generation event store
Read original