slack

From Custom to Open: Scalable Network Probing and HTTP/3 Readiness with Prometheus (opens in new tab)

Slack needed better client-side observability while migrating edge services to HTTP/3, which uses QUIC over UDP rather than TCP. Existing SaaS tools and Prometheus Blackbox Exporter could not probe HTTP/3 endpoints, so an intern added QUIC support using Go’s quic-go library and open-sourced it. The result unified HTTP/1.1, HTTP/2, and HTTP/3 monitoring while making the capability available to the broader Prometheus community.

Limitations of Legacy Monitoring

  • Slack used a mix of commercial monitoring services and internal tools for network measurements.
  • HTTP/3 introduced a major observability gap because it runs over QUIC/UDP.
  • Existing SaaS solutions lacked built-in HTTP/3 probing.
  • Prometheus Blackbox Exporter had no native QUIC support.
  • Without probing at scale, Slack could not reliably measure round-trip times, detect regressions to HTTP/2, or monitor hundreds of thousands of HTTP/3 endpoints.

Adding QUIC Support to Blackbox Exporter

  • Intern Sebastian Feliciano selected quic-go because of its adoption and first-class Go HTTP client support.

  • The implementation used an http3.Transport with TLS and QUIC configuration:

    http3Transport := &http3.Transport{
        TLSClientConfig: tlsConfig,
        QUICConfig:      &quic.Config{},
    }
    
  • The new transport was attached to a standard Go http.Client.

  • The implementation preserved Blackbox Exporter’s existing configuration and composability patterns.

  • Sebastian open-sourced the feature and eventually got it accepted upstream.

In-House Integration and Operational Benefits

  • Because upstream review could take longer than the internship timeline, Slack built an internal system around the new functionality.
  • Grafana now provides a unified view of HTTP/1.1, HTTP/2, and HTTP/3 metrics.
  • Operators can compare protocol performance and correlate it with other telemetry.
  • Improved visibility supports more accurate alerts and faster debugging of HTTP/3 issues.

Future Enhancements

  • SNI routing tests: Verify that shared edge infrastructure routes hostnames to the correct backend and presents the correct TLS certificate.
  • End-to-end path visualization: Map network hops between monitoring agents and endpoints to identify latency spikes or packet loss more precisely.

Broader Lessons

  • Observability should be established before a major protocol or infrastructure migration.
  • Filling gaps through open source can benefit both the organization and the wider engineering community.
  • Supporting emerging protocols such as QUIC early helps future-proof monitoring systems.

Slack recommends trying the new QUIC functionality in Prometheus Blackbox Exporter and contributing to its continued development.