datadog3 min read

Curated summary

Scaling down to speed up: How we improved efficiency of live process metrics by 100x

Read original(opens in new tab)

Datadog redesigned its real-time Processes and Containers pipeline to avoid collecting high-frequency metrics that users never see. By limiting 2-second collection to hosts actively viewed and using standard 10-second data for sorting, the company reduced real-time traffic by over 100x, cut infrastructure costs by 98%, and lowered Agent resource usage. The approach also improved scalability without sacrificing the live investigation experience.

Original Real-Time Collection Model

  • Datadog Agents normally collect process and container metrics every 10 seconds.
  • When a user opened a live Processes or Containers view, all hosts in that tenant switched to 2-second collection.
  • This supported near-real-time monitoring similar to htop, but across distributed infrastructure.
  • As tenants grew, the pipeline had to process millions of processes per second, even though users typically viewed only around 50 processes or containers.
  • Live sorting required keeping all tenant data in memory on a single server, limiting horizontal scaling and forcing vertical scaling.

Refocusing on User-Visible Data

  • Most collected metrics were never displayed to users.
  • Datadog determined that real-time collection only needed to be enabled for hosts running the processes or containers currently in view—up to roughly 50 hosts per user.
  • Internal telemetry suggested this could reduce traffic by more than 100x.
  • This required tracking active host subscriptions and updating them as users navigated the product.
  • Because sorting occurred every 10 seconds, it did not need 2-second data. Datadog switched live views to use the existing 10-second metrics, aligning live and historical sorting logic.

Host Subscription Filtering

  • A proof of concept added host subscriptions to the live data servers.
  • Servers filtered Kafka payloads and discarded data for hosts without active subscriptions.
  • This immediately reduced:
    • Memory usage by 85%
    • CPU usage by 33%
  • The improvement came from storing fewer live metrics and processing fewer incoming payloads.
  • The prototype confirmed that filtering preserved product behavior while simplifying sorting.

Moving Filtering Earlier in the Pipeline

  • Late filtering improved live data servers but still left unnecessary work for the rest of the system and customer-side Datadog Agents.
  • Datadog therefore planned to propagate subscription state to the intake service.
  • Live data servers publish users’ active host sets over Kafka once per second.
  • The intake service consumes this information and decides which hosts should activate 2-second process and container metric collection.
  • This allows real-time collection to be restricted to hosts users are actively investigating while maintaining responsive live views.

Datadog’s redesign demonstrates that real-time systems scale more effectively when they prioritize data users can actually see. Filtering at intake, limiting high-frequency collection to subscribed hosts, and reusing standard-resolution data for sorting provide a simpler and more economical architecture without eliminating live functionality.

Continue with another curated summary.