Curated summary
Analyzing Incident Causes with Natural Language in Grafana: Developing an LLM Agent-Based SRELens
SRELens is a Grafana-based natural-language observability assistant created by LY Corporation’s Home SRE team. It connects metrics, logs, traces, and profiles so engineers can investigate incidents without switching between tools or manually transferring context. The project’s central conclusion is that production reliability depends less on natural-language querying itself and more on controlling the LLM’s tools, prompts, permissions, cost, and failure behavior through backend code and policy.
The Observability Analysis Problem
- Incident investigation traditionally requires moving among:
- Grafana or IMON for metrics
- LaaS or IU for logs
- IMON Trace or Tempo for traces
- A separate profiling system
- Engineers must manually connect:
- Error-rate increases
- Error messages
- Trace IDs and slow requests
- Relevant time ranges, services, and labels
- This context switching is especially costly during outages.
- The team first consolidated data with a self-hosted LGTM-P stack:
- Mimir for metrics
- Loki for logs
- Tempo for traces
- Pyroscope for profiles
- OpenTelemetry Collector as the ingestion layer
- Centralizing the data helped, but engineers still needed to know the correct datasource, labels, query syntax, and relationships between signals.
Why an Existing Open-Source PoC Was Not Enough
The team initially evaluated an open-source Grafana LLM plugin, but identified several production limitations:
- It could not reliably propagate Grafana-authenticated user context for chat history, permissions, and usage limits.
- System prompts could not be controlled strongly enough to enforce organizational policies.
- Short tool-call limits interrupted multi-step investigations.
- Datasource-specific naming differences often produced empty results:
- Metrics might use
service_name - Tempo might require
resource.service.name - Loki might require JSON parsing or structured metadata filters
- Metrics might use
- Modifying and deploying the solution internally raised operational and licensing concerns.
The PoC showed that the key requirement was not merely asking questions in natural language, but retaining control over how the agent operates.
SRELens Architecture
- SRELens runs as a Grafana application plugin.
- The frontend provides the chat interface.
- The backend handles:
- LLM requests
- Tool orchestration
- Prompt composition
- Usage and quota enforcement
- Observability queries are executed through an MCP gateway.
- A
CompositeClientcombines:- Upstream FlavaMCP observability tools
- Local Grafana tools such as
find_grafana_panelandrender_grafana_panel
- The backend is an orchestration and policy layer, not just a proxy.
Three-Layer System Prompt Design
Base System Prompt
Defines organization-wide behavior and safety rules, including:
- Tool-call ordering
- Safe handling of dashboard creation, modification, and deletion
- Fallback behavior for empty results
- Re-querying with aggregation when results are truncated
- Response structure and evidence requirements
Only administrators can change this layer.
Datasource Fragment
Encodes environment-specific operational knowledge in YAML:
- Preferred Mimir, Loki, and Tempo datasource UIDs
- Candidate service-name labels
- Loki parsing and filtering rules
This prevents the agent from wasting tool-call rounds discovering basic datasource conventions.
User Prompt
Stores personal or team-specific context in Redis, such as:
- Owned services
- Preferred response formats
- Frequently used dashboards
User preferences are added as context but cannot override organizational safety policies.
Backend Tool Orchestration and Guardrails
The backend exclusively assembles system prompts and runs the agent loop:
- Send the user’s question to the LLM.
- Execute requested MCP or local tools.
- Return tool results to the LLM.
- Repeat until a final answer is produced.
Safety and reliability controls include:
- A default maximum of 10 tool-call rounds
- Duplicate-call prevention using call hashes
- A default retry limit of two attempts per tool
- Per-tool result-size limits
- Trimming older tool results when the request history becomes too large
- Preserving
tool_call_idrelationships when trimming history - Hints that encourage changing labels, time ranges, or datasources after empty results
These safeguards reduce dependence on the LLM making perfect decisions.
Usage Limits and Degraded Operation
- Per-user daily token quotas
- Per-user requests-per-minute limits
- HTTP 429 responses after limits are exceeded
- Post-response accounting based on actual prompt and completion tokens returned by OpenAI
- Daily quota reset at midnight in the Asia/Seoul timezone
- Redis stores conversation history, user prompts, and quotas.
- If Redis is unavailable, personalization and history are reduced, but a single chat request can still proceed.
Incident Analysis Scenario
In one beta service, SRELens was asked to investigate an error spike between 09:50 and 10:05.
- Instead of separately searching alerts, logs, and traces, the agent examined the relevant dashboard and observability data together.
- It narrowed the incident to a surge in
CopyMediarequests. - The analysis was intended to connect the request pattern with the underlying errors and supporting telemetry, demonstrating how SRELens can move from an aggregate error spike toward a specific API-level cause.
SRELens demonstrates that an LLM can accelerate incident analysis when it is grounded in an integrated observability stack and constrained by explicit backend policies. For production use, organizations should treat prompt control, tool orchestration, permissions, quotas, retries, and failure handling as core system components rather than leaving them entirely to the model.
Related reading
Continue with another curated summary.
How we used AI agents to migrate GitLab rate limiting
Read originalStress Testing Know-How for Messaging Servers and How AI Lightened the Load
Read originalHow to build CI/CD observability at scale
Read originalFrom Custom to Open: Scalable Network Probing and HTTP/3 Readiness with Prometheus
Read original