fastapi

2 posts

netflix

In-House LLM Serving at Netflix (opens in new tab)

Netflix built an in-house LLM serving platform within its existing production ML infrastructure rather than creating a separate ML stack. The platform combines a JVM-based serving layer, NVIDIA Triton, GPU-backed Model Scoring Service, and an OpenAI-compatible HTTP frontend. Its main design choices—vLLM, model packaging, API compatibility, and deployment strategy—prioritize operational flexibility and seamless movement from hosted models to self-hosted ones, while production exposed versioning and compatibility risks. ## Architecture and Serving Model - Netflix’s unified JVM serving system handles routing, A/B testing, feature retrieval, inference, post-processing, and logging. - Callers access models through: - A gRPC path integrated with the existing serving system. - A direct HTTP path for newer LLM applications. - Small CPU models run in-process to avoid remote-call overhead. - Larger GPU models run through Model Scoring Service (MSS), which supports XGBoost, TensorFlow, PyTorch, and LLMs. - NVIDIA Triton manages model loading, batching, and GPU scheduling. - A Java control plane provides deployment, versioning, health checks, autoscaling, and multi-region rollout. ## Choosing vLLM as the Standard Engine - Netflix originally used TensorRT-LLM, but re-evaluated its choice as open-source engines improved and workloads diversified. - vLLM was selected based on operational fit rather than benchmark performance alone: - Supports custom model architectures without lengthy compilation. - Provides hooks for custom decoding and constraint logic. - Is easier to debug than earlier compiled-engine workflows. - Is familiar to many researchers, reducing the research-to-production transition cost. - The workload includes embeddings, prefill-only inference, autoregressive decoding, and custom per-step decoding constraints. ## Triton Integration and Model Packaging - Triton offers both a Python backend and a dedicated vLLM backend. - The Python backend requires explicit input and output tensor definitions, coupling packaged artifacts to frontend changes. - The vLLM backend uses a JSON configuration pointing to model weights and tokenizers, generating tensor specifications dynamically. - Netflix considers the vLLM backend the preferred default because models and frontends can evolve independently. - Production revealed two limitations: - Triton and vLLM must be version-pinned because incompatible APIs can prevent the backend from loading entirely. - Models requiring custom preprocessing, postprocessing, tokenization, or ensemble execution still need Triton’s Python backend. ## OpenAI-Compatible HTTP Frontend - Netflix keeps LLMs compatible with the same internal gRPC model-serving interface used by other model types. - It also exposes an OpenAI-compatible API because that interface is widely supported by inference engines, orchestration tools, evaluation systems, and client libraries. - This makes replacing a hosted model with a fine-tuned self-hosted model largely transparent to callers. - The implementation uses Triton’s OpenAI-compatible frontend, FastAPI, and a `TritonLLMEngine` that translates requests into Triton inference calls. - KServe HTTP and gRPC frontends remain available for the Java control plane. - Netflix found that Triton’s frontend silently discarded the `response_format` parameter, meaning JSON requests could reach vLLM without guided decoding and produce malformed output. - The team patched the frontend to translate `response_format` into vLLM guided-decoding parameters. ## Deployment and Rollout Strategies - GPU services require longer startup times than CPU services, and model versions may change input/output schemas. - Netflix supports Red-Black deployment: - Runs the new version alongside the old one. - Performs health checks before shifting traffic. - Gradually scales up the new version while scaling down the old one. - Supports atomic rollback if deployment fails. - Red-Black deployment works well when the model interface remains stable. - Production exposed a schema-coordination problem: if a new model changes tensor dimensions or other I/O requirements, upstream callers may send old requests to the new model during the migration window, causing failures. - The post introduces a Versioned strategy as a solution, but the provided text ends before explaining its implementation. Netflix’s experience suggests that successful in-house LLM serving depends as much on compatibility and deployment mechanics as on raw inference speed. A practical platform should standardize on an extensible engine such as vLLM, preserve ecosystem-compatible APIs, tightly control engine versions, retain escape hatches for custom models, and explicitly coordinate model-schema changes during rollout.

kakao

Bringing a Voice AI Model to Production: The Journey of Optimizing Kanana-O Serving (opens in new tab)

Kanana-O is a multimodal model that understands text, images, and audio, then responds with text and speech. Deploying it for real-time voice conversations required solving problems that do not arise during model training, including low first-response latency, concurrent users, streaming across multiple models, and uneven GPU memory demands. Kakao built the specialized Kanana-Omni Server, achieving 1.6× the throughput of vLLM-Omni at 64 concurrent users. ## Kanana-O’s Three-Stage Architecture - **Thinker** processes multimodal inputs and generates text. - **Talker** converts Thinker’s text embeddings into sequential speech tokens. - **VoiceBox** combines speech tokens into audible audio waveforms. - In production, these components must operate concurrently rather than sequentially to deliver audio within hundreds of milliseconds. ## Why a Specialized Serving Server Was Needed - Thinker passes hidden-state embeddings directly to Talker rather than ordinary token IDs. - These high-dimensional tensors must be transferred continuously, making serialization or CPU copies too expensive. - Talker produces speech tokens step by step, while VoiceBox waits for enough tokens to form larger audio chunks. - Talker also combines speaker embeddings, Thinker outputs, and its own accumulated audio embeddings, creating an input structure unlike standard autoregressive decoding. - These constraints made a custom server more suitable than general-purpose frameworks. ## Zero-Copy Data Transfer - The server preallocates shared-memory blocks during startup. - Thinker writes tensors into an available block, while Talker receives only metadata such as the block identifier and byte size. - This avoids repeated allocation, copying, and serialization. - For GPU tensors on the same node, CUDA IPC transfers data directly between GPU processes, avoiding Device→Host→Device movement. ## Cascaded Streaming Pipeline - Thinker, Talker, and VoiceBox run as overlapping asynchronous stages. - Thinker can send its first output chunk while Talker processes earlier chunks and VoiceBox synthesizes audio from still earlier ones. - Talker buffers speech tokens until VoiceBox has enough data to create an audio chunk. - This pipelining significantly reduces the time before the user hears the first response. ## Process Isolation and Fault Containment - Thinker and Talker each run their own vLLM engine in separate processes. - This avoids conflicts between CUDA contexts, model memory, KV caches, and schedulers. - Processes are started with `spawn` rather than `fork`, preventing inherited CUDA state from causing corruption. - If one component fails, such as Thinker running out of memory, the other components and the API server can continue operating and be restarted independently. ## Continuous Batching with vLLM - Manually batching requests is difficult because multimodal inputs and accumulated Talker embeddings vary in size. - The server submits requests rapidly and delegates batch construction to vLLM’s continuous-batching scheduler. - Each request runs as an independent asynchronous generation task. - vLLM combines requests internally during forward passes, while request IDs ensure each task receives only its own streamed output. - This improves GPU utilization without requiring custom synchronization and padding logic. ## Single FastAPI Worker and Asynchronous Execution - Multiple Uvicorn workers would load separate copies of the vLLM engines, multiplying GPU memory usage and model-loading costs. - Therefore, the server uses `workers=1`. - Since a blocking operation would otherwise stall every connected user, the entire request path—from the API endpoint through final audio generation—is designed around `async`/`await`. - Keeping the pipeline non-blocking allows one worker to accept and progress many concurrent requests. Kakao’s main recommendation is to design serving infrastructure around the model’s actual dataflow rather than forcing it into a generic framework. For complex multimodal pipelines, zero-copy transfers, asynchronous cascaded streaming, process isolation, and engine-level continuous batching can be more important than simply scaling API workers.