kakao

Bringing a Voice AI Model to Production: The Journey of Optimizing Kanana-O Serving (opens in new tab)

Kanana-O is a multimodal model that understands text, images, and audio, then responds with text and speech. Deploying it for real-time voice conversations required solving problems that do not arise during model training, including low first-response latency, concurrent users, streaming across multiple models, and uneven GPU memory demands. Kakao built the specialized Kanana-Omni Server, achieving 1.6× the throughput of vLLM-Omni at 64 concurrent users.

Kanana-O’s Three-Stage Architecture

  • Thinker processes multimodal inputs and generates text.
  • Talker converts Thinker’s text embeddings into sequential speech tokens.
  • VoiceBox combines speech tokens into audible audio waveforms.
  • In production, these components must operate concurrently rather than sequentially to deliver audio within hundreds of milliseconds.

Why a Specialized Serving Server Was Needed

  • Thinker passes hidden-state embeddings directly to Talker rather than ordinary token IDs.
  • These high-dimensional tensors must be transferred continuously, making serialization or CPU copies too expensive.
  • Talker produces speech tokens step by step, while VoiceBox waits for enough tokens to form larger audio chunks.
  • Talker also combines speaker embeddings, Thinker outputs, and its own accumulated audio embeddings, creating an input structure unlike standard autoregressive decoding.
  • These constraints made a custom server more suitable than general-purpose frameworks.

Zero-Copy Data Transfer

  • The server preallocates shared-memory blocks during startup.
  • Thinker writes tensors into an available block, while Talker receives only metadata such as the block identifier and byte size.
  • This avoids repeated allocation, copying, and serialization.
  • For GPU tensors on the same node, CUDA IPC transfers data directly between GPU processes, avoiding Device→Host→Device movement.

Cascaded Streaming Pipeline

  • Thinker, Talker, and VoiceBox run as overlapping asynchronous stages.
  • Thinker can send its first output chunk while Talker processes earlier chunks and VoiceBox synthesizes audio from still earlier ones.
  • Talker buffers speech tokens until VoiceBox has enough data to create an audio chunk.
  • This pipelining significantly reduces the time before the user hears the first response.

Process Isolation and Fault Containment

  • Thinker and Talker each run their own vLLM engine in separate processes.
  • This avoids conflicts between CUDA contexts, model memory, KV caches, and schedulers.
  • Processes are started with spawn rather than fork, preventing inherited CUDA state from causing corruption.
  • If one component fails, such as Thinker running out of memory, the other components and the API server can continue operating and be restarted independently.

Continuous Batching with vLLM

  • Manually batching requests is difficult because multimodal inputs and accumulated Talker embeddings vary in size.
  • The server submits requests rapidly and delegates batch construction to vLLM’s continuous-batching scheduler.
  • Each request runs as an independent asynchronous generation task.
  • vLLM combines requests internally during forward passes, while request IDs ensure each task receives only its own streamed output.
  • This improves GPU utilization without requiring custom synchronization and padding logic.

Single FastAPI Worker and Asynchronous Execution

  • Multiple Uvicorn workers would load separate copies of the vLLM engines, multiplying GPU memory usage and model-loading costs.
  • Therefore, the server uses workers=1.
  • Since a blocking operation would otherwise stall every connected user, the entire request path—from the API endpoint through final audio generation—is designed around async/await.
  • Keeping the pipeline non-blocking allows one worker to accept and progress many concurrent requests.

Kakao’s main recommendation is to design serving infrastructure around the model’s actual dataflow rather than forcing it into a generic framework. For complex multimodal pipelines, zero-copy transfers, asynchronous cascaded streaming, process isolation, and engine-level continuous batching can be more important than simply scaling API workers.