Curated summary
How we built a real-world evaluation platform for autonomous SRE agents at scale
Bits AI SRE improved in isolated scenarios but lacked a way to detect regressions across the broader range of production incidents. The team found that tool-level tests and live replays could not capture failures caused by multi-step reasoning or changing telemetry. They built a replayable evaluation platform combining realistic investigation labels, scalable orchestration, and longitudinal performance tracking.
Subtle Regressions from Well-Intentioned Features
- Adding a monitor’s service name to the agent’s initial context improved some internal investigations.
- Across broader scenarios, it introduced irrelevant signals that confused the agent and degraded unrelated investigations.
- Because there was no representative evaluation set, the team could not measure the change’s wider impact before internal misses exposed it.
- The incident demonstrated the need to evaluate every change across diverse investigation types.
Limits of Tool Tests and Live Replay
- Testing tools individually failed to capture errors caused by incorrect interactions between valid tool outputs.
- Live investigation replay was difficult to scale because:
- Results were not consistently aggregated.
- Production environments changed.
- Telemetry expired, making investigations unreplayable.
- Standard evaluation frameworks assumed clean inputs and static datasets, unlike agents operating over production telemetry.
- The team needed controlled, offline replay of realistic end-to-end investigations.
Evaluation Labels and World Snapshots
- Each label represents one production-style investigation.
- It contains:
- Ground truth: the issue’s actual root cause.
- World snapshot: the queries and signals available when the issue occurred.
- The agent is never shown the root cause directly; it must reason from the preserved signals.
- Labels must cover varied technologies and failure modes, including:
- Kubernetes pod failures
- Kafka lag
- Bad-code deployments
- Complex multi-service business failures
- A narrow or overly clean dataset would make performance appear better than it really is.
Orchestrating Evaluations at Scale
- The platform runs Bits against labels, scores the outcomes, and tracks quality over time.
- It supports comparisons across:
- Investigation categories
- Model variants
- Configuration versions
- Evaluation runs
- The architecture consists of a shared label set, an orchestration layer, and reporting infrastructure.
- This allows teams to determine whether improvements in one domain, such as Kafka, regress another, such as Kubernetes.
Scaling Label Creation
- The team initially created labels manually from Datadog alerts.
- Manual labeling provided early coverage but consumed engineering time and remained far from representative.
- They embedded label generation into Bits itself:
- Customer feedback and investigation data are used to derive root causes.
- Relevant queries are preserved as the world snapshot.
- Each user interaction becomes a potential evaluation case.
- This increased label creation rates by an order of magnitude and allowed coverage to grow with product usage.
Agent-Assisted Validation
- Early labels required extensive human review, especially when feedback was ambiguous or reconstructed signals were uncertain.
- As ingestion grew, manual review became a bottleneck.
- Bits was then used to assist with validation by aggregating related signals, identifying relationships, and resolving ambiguous feedback before human review.
Practical Conclusion
Reliable agent improvement requires more than testing individual tools or replaying live incidents. A representative, production-derived label set combined with reproducible end-to-end evaluations makes regressions visible and enables safer iteration.
Related reading
Continue with another curated summary.
When upserts don't update but still write: Debugging Postgres performance at scale
Read originalWhen an AI agent came knocking: Catching malicious contributions in Datadog’s open source repos
Read originalDesigning MCP tools for agents: Lessons from building Datadog's MCP server | Datadog
Read originalHow we reduced the size of our Agent Go binaries by up to 77% | Datadog
Read original