latency-optimization

2 posts

cloudflare

Agents Week: network performance update (opens in new tab)

Cloudflare reports that it became the fastest network in 60% of the world’s 1,000 largest networks by December 2025, up from 40% during Birthday Week 2025. The improvement came from both expanding its global points of presence and optimizing connection-handling software. Cloudflare says it is continuing to target the remaining networks where competitors still lead. ### Measuring Network Performance - Cloudflare analyzes the 1,000 largest networks by estimated population, using APNIC data. - It measures TCP connection time—the time required to complete a TCP handshake—as a practical indicator of users’ perceived Internet speed. - Rankings use the **trimean**, a weighted average of the 25th percentile, median, and 75th percentile, reducing the influence of outliers. - Data comes from Real User Measurements: browsers loading Cloudflare error pages silently download small files from Cloudflare, Amazon CloudFront, Google, Fastly, and Akamai under real network conditions. ### Expanding Points of Presence - New locations in Constantine, Algeria; Malang, Indonesia; and Wroclaw, Poland brought Cloudflare physically closer to users. - In Wroclaw, free-user average round-trip time fell from 19 ms to 12 ms, a 40% improvement. - In Malang, Enterprise traffic improved from 39 ms to 37 ms, a 5% reduction. - However, new locations alone did not account for the increase from 40% to 60% of networks. ### Improving Connection Handling - Cloudflare optimized the software responsible for connection establishment, SSL/TLS termination, traffic management, and proxying. - HTTP/3 adoption and changes to congestion-window management reduced processing time. - Improvements in CPU and memory efficiency allow the global network to handle connections more effectively. - Cloudflare compares this to improving both the efficiency of highway toll booths and the routing of traffic between them. ### Results by December 2025 - Cloudflare was the fastest provider in 60% of the largest networks. - Between September and December 2025, it became fastest in: - 40 additional countries - 261 additional networks - 54 additional U.S. autonomous systems (ASNs) - During December, Cloudflare was on average 6 ms faster than the next-fastest provider. Cloudflare’s conclusion is that continued gains in network reach and software efficiency can produce measurable improvements for users. It plans to focus on the remaining networks where it is narrowly behind competitors, with the long-term goal of being fastest worldwide.

pinterest

Beyond Two Towers: Re-architecting the Serving Stack for Next-Gen Ads Lightweight Ranking Models… (opens in new tab)

Two-Tower models make retrieval and lightweight ranking highly efficient by scoring user and item embeddings with a dot product, but they cannot represent rich user-item interactions or deep feature crossings. This post describes an ads-serving redesign that introduces general-purpose GPU models while preserving end-to-end latency. The main strategy is to reduce data movement, move filtering logic onto the GPU, and optimize inference from an initial 4-second p90 latency to about 20 milliseconds. ## Why Move Beyond Two-Tower Models - Two-Tower architectures independently encode users and items, enabling fast scoring across millions of candidates. - Their decoupled structure limits: - User-item interaction features - Target attention - Early feature crossing - Deep architectures requiring simultaneous access to user and candidate data - More expressive models require GPU-based general-purpose inference rather than specialized dot-product or ANN retrieval. - The existing retrieval stack was not designed to transfer large candidate and feature sets to a GPU, creating a major latency challenge. ## Restructuring the Serving Funnel The traditional funnel consisted of: - Feature expansion for thousands of candidates - Retrieval and Two-Tower lightweight ranking - Heavy ranking and auction processing for the top documents Adding GPU inference directly to this flow would require fetching, serializing, transferring, and returning features for tens of thousands of documents. The authors therefore redesigned the entire early-stage serving pipeline instead of optimizing the model alone. ## Segmenting the Inventory for Feature Fetching Feature retrieval was a major latency source, often taking longer than model inference for workloads ranging from 10,000 to 100,000 documents. - **High-value inventory:** Roughly 1 million documents responsible for a substantial share of revenue have their features embedded in the PyTorch model as registered buffers. - Features become part of the model state, similar to weights. - They remain in GPU high-bandwidth memory. - Requests avoid remote feature-service calls and host-to-device transfers. - The model file must be periodically updated to refresh features. - Future work may include GPU-based caching. - **Long-tail inventory:** The remaining roughly 1 billion documents use a high-performance key-value store with in-host caching. - The post focuses on the first strategy, which is already running in production. ## Moving Business Logic onto the GPU Previously, the model returned scores for approximately 100,000 candidates, while CPU-side code handled utility calculation, filtering, diversity, deduplication, and top-k selection. - The new PyTorch model performs these operations directly: - Combines pCTR, pCVR, bid, and other signals into utility scores. - Applies diversity and filtering rules. - Performs top-k selection. - The GPU returns only the final winners—typically around 1,000 documents—instead of all candidate scores. - This reduces device-to-host data transfer and takes advantage of GPU parallelism. - The approach works because lightweight-ranking business rules are sufficiently simple to express with tensor operations. ## Reducing GPU Inference Latency Initial GPU inference measured roughly 4,000 ms at p90, far too slow for real-time serving. Several systems optimizations reduced this to approximately 20 ms: - **Multiple CUDA streams:** Separate streams for workers allow host-to-device transfers, computation, and device-to-host transfers to overlap. - **Worker alignment:** Worker threads are matched and pinned to physical CPU cores to reduce context switching and lock contention. - **Kernel fusion:** Triton kernels combine operations such as linear layers and activations, reducing memory traffic. - **BF16 computation:** Brain Floating Point 16 lowers memory usage and accelerates arithmetic compared with FP32. - **Profiling tools:** PyTorch Profiler and NVIDIA Nsight Systems were used to identify bottlenecks. ## Practical Recommendation Deploying more expressive ranking models requires rethinking the serving architecture around data movement and execution placement. Embedding frequently used features, executing business logic on the GPU, and applying low-level CUDA and kernel optimizations can make complex neural ranking feasible without increasing end-to-end latency.