KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure (opens in new tab)
KernelEvolve is Meta’s agentic system for automating the creation and optimization of low-level AI kernels across diverse hardware. It treats kernel tuning as a search problem rather than one-shot code generation, evaluating hundreds of alternatives with profiling and diagnostics. The system reduces optimization work from weeks to hours and has delivered over 60% higher inference throughput for an Ads model on NVIDIA GPUs and over 25% higher training throughput on Meta’s MTIA chips.
Kernel Optimization at Meta
- AI models rely on optimized kernels that translate high-level operations into hardware-specific instructions.
- Meta runs models across NVIDIA GPUs, AMD GPUs, custom MTIA accelerators, and CPUs.
- Production workloads require many custom operators beyond standard GEMMs and convolutions available in vendor libraries.
- Kernels must be developed and tuned for each combination of:
- Hardware type and generation
- Model architecture
- Operator type
The Challenge of Hardware Heterogeneity
- NVIDIA, AMD, MTIA, and CPU platforms differ in:
- Memory architectures and hierarchies
- Instruction sets
- Execution models
- Supported numeric data types
- A kernel optimized for one platform may perform poorly or fail on another.
- Hardware generations also require new optimization strategies. Meta’s MTIA roadmap includes four generations, from MTIA 300 through MTIA 500, in two years.
- Manual tuning by kernel specialists cannot keep pace with these changes.
Increasing Model and Operator Complexity
- Meta’s recommendation systems have evolved from embedding-based models to sequence models with attention, GEM, and LLM-scale models such as Meta Adaptive Ranking Model.
- Each new model generation introduces operators that earlier systems did not require.
- Multiple model families may be involved in a single ads-serving request.
- As model architectures and operator inventories grow, the number of kernel configurations expands rapidly into the thousands.
How KernelEvolve Works
- KernelEvolve generates candidate implementations in languages and DSLs including:
- Triton, Cute DSL, and FlyDSL
- CUDA, HIP, and MTIA C++
- A dedicated job harness compiles, runs, profiles, and evaluates each candidate.
- Performance results, correctness checks, and diagnostic information are fed back to the LLM.
- The system continuously searches through hundreds of alternatives instead of stopping at the first plausible implementation.
- Its automated workflow includes profiling, optimization, testing, and cross-hardware debugging.
Results and Broader Impact
- KernelEvolve improved Andromeda Ads inference throughput by more than 60% on NVIDIA GPUs.
- It improved training throughput for an ads model by more than 25% on Meta’s MTIA silicon.
- The system operates across both public and proprietary hardware.
- In production, it optimizes code supporting trillions of daily inference requests.
- By automating kernel development, Meta can enable new hardware and adapt to changing model architectures with substantially less engineering effort.
KernelEvolve turns kernel development from a manual, expert-driven bottleneck into a continuous automated process. Its search-based approach is particularly valuable as Meta’s hardware portfolio and model architectures continue to diversify.