Curated summary
Making fetch happen: Building a general-purpose query and render scheduler
Datadog rebuilt its dashboard scheduler to improve responsiveness while distributing network and rendering work more efficiently. The legacy system helped, but had grown into a complex set of roughly 20 interdependent heuristics that were difficult to maintain and poorly separated query scheduling from rendering. A simpler, general-purpose approach reduced request spikes, improved fetching performance, and created a foundation for browser-aware task scheduling.
Limitations of the Original Scheduler
- A periodic updater determined when widgets should request fresh data based on factors such as time range and browser focus.
- Query tasks for visible widgets ran immediately; offscreen queries were delayed using heuristics such as pending-query counts and historical fetch durations.
- Render tasks similarly prioritized visible widgets and delayed offscreen work.
- The system improved performance over an unscheduled baseline by reducing main-thread work and flattening query traffic.
- Over time, it accumulated around 20 parameters and interlinked rules.
- Query and render concerns were mixed together:
- Queries could be delayed because too many render tasks were pending.
- Renders could be delayed based on data size even when browser resources were available.
- The dashboard-specific implementation could not easily be reused across Datadog’s increasingly generalized widget framework.
A General-Purpose Scheduling Strategy
- Datadog separated query scheduling from render scheduling so each could be developed, tested, and rolled out independently.
- The team evaluated existing heuristics across dashboards of different sizes and under different browser conditions.
- Several rules were removed without harming performance:
- Unfocused or occluded tabs did not need special delays because the periodic updater and browser already throttle them.
- The redesign aimed to preserve two goals:
- Keep query execution distributed over time.
- Prioritize widgets visible to the user.
- The new system was progressively deployed, first to dashboards and then to the shared data-fetching framework used across Datadog.
Simpler Query Scheduling
The new query algorithm uses a small set of straightforward rules:
- Fetches for visible widgets run immediately.
- Non-visible queries are ranked and executed in fixed time windows, subject to a task limit.
- Query execution pauses when the number of pending fetches becomes too high.
- The chosen configuration uses:
- A 2,000-millisecond time window.
- A maximum of 10 tasks per window.
- FIFO-style ranking for offscreen queries, favoring earlier requests.
- The scheduler uses only about six parameters instead of the legacy system’s roughly 20.
- The simplified algorithm produced a better task distribution than the old scheduler.
- “429 Too many requests” errors dropped significantly, reducing retries and helping data arrive sooner.
Browser-Aware Render Scheduling
- The old render scheduler did not account for the browser’s available CPU and memory resources.
- Datadog adopted the Browser Scheduling API to create prioritized tasks that the browser can schedule natively.
- Tasks can receive priorities such as:
user-blockinguser-visiblebackground
- A
TaskControllerassigns a priority signal to scheduled work. - Priorities can later be changed for all tasks controlled by the same controller, and tasks can be aborted.
- The API was supported in Chromium and Firefox Nightly, with a polyfill for other browsers.
Datadog’s experience suggests that performance schedulers benefit from simple, independently testable rules: prioritize visible work, smooth network activity, and let the browser manage expensive rendering when possible.
Related reading
Continue with another curated summary.
Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale
Read originalHow we migrated a live routing system using AI-assisted refactoring
Read originalSteganography at scale: Embedding share URLs in Datadog widget screenshots
Read originalWhen upserts don't update but still write: Debugging Postgres performance at scale
Read original