Curated summary
Powering the agents: Workers AI now runs large models, starting with Kimi K2.5
Database DesignLarge Language ModelsKubernetesCloudflareOpen SourceWorkers AiAi ModelsInference Engine
Cloudflare is expanding Workers AI beyond smaller models by adding Moonshot AI’s Kimi K2.5, a frontier open-source model designed for agentic workloads. With a 256k context window, tool calling, vision, and structured outputs, Kimi can power an agent’s full lifecycle directly on Cloudflare’s platform. Cloudflare argues that its price-performance makes open-source models essential as personal and enterprise agents dramatically increase inference demand.
Kimi K2.5’s Price-Performance Advantage
- Cloudflare uses Kimi internally for:
- Agentic coding through OpenCode
- Automated code review via the Bonk public code review agent
- Security analysis of Cloudflare codebases
- A security-review agent processes more than 7 billion tokens daily and has found over 15 confirmed issues in one codebase.
- Compared with a mid-tier proprietary model, switching to Kimi reduced the estimated cost of this workload by 77%, from roughly $2.4 million annually.
- As employees increasingly run multiple agents continuously, inference costs become a major barrier to scaling.
- Cloudflare positions open-source, frontier-quality models as a more economical alternative to proprietary systems.
Serving Large Models on Workers AI
- Supporting Kimi required upgrades to Workers AI’s inference stack, which historically focused on smaller models.
- Cloudflare uses its proprietary Infire inference engine and custom kernels to improve:
- Model performance
- GPU utilization
- Throughput
- The platform applies advanced serving strategies such as:
- Data, tensor, and expert parallelization
- Disaggregated prefill, separating input processing from generation across machines
- Workers AI handles these infrastructure optimizations so developers do not need specialized machine learning, DevOps, or reliability engineering expertise.
Prefix Caching for Agent Workloads
- Agents frequently resend large prompts containing:
- System instructions
- Tool definitions
- MCP server tools
- Conversation history
- Entire codebases
- Prefix caching avoids reprocessing unchanged input tokens during multi-turn interactions.
- This reduces prefill work, improving:
- Time to First Token (TTFT)
- Tokens Per Second (TPS)
- Overall inference cost
- Workers AI now exposes cached tokens as a usage metric and charges less for them than regular input tokens.
- Cloudflare has also introduced techniques to improve cache hit rates.
Session Affinity
- Workers AI provides an
x-session-affinityheader to route requests from the same session or agent to the same model instance. - Keeping requests on the same instance increases prefix-cache reuse.
- Higher cache hit rates lead to faster responses, greater throughput, and lower costs.
- Clients should provide a unique session or agent identifier with the header.
Cloudflare’s recommendation is to use Workers AI when building agents that need frontier-level reasoning without the cost and operational burden of proprietary models or self-hosted infrastructure.
Related reading
Continue with another curated summary.