BE
Berkeley BAIR
7/29/2026
From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
Short summary
Berkeley researchers extended K-Search, an evolutionary kernel optimization framework, with a CUDA-to-MLX translation layer that transfers decades of CUDA kernel expertise to Apple Silicon. The approach achieves 0.97x speedup versus native MLX Attention and up to 20x prefill speedup on Mamba SSM kernels. The method generalizes to any ecosystem where CUDA optimization knowledge is transferable.
- •K-Search uses LLM-guided evolutionary search to optimize GPU kernels across hardware
- •CUDA-to-MLX translation layer adapts existing CUDA expertise rather than rebuilding from scratch
- •Near-expert performance on Apple Silicon: 0.97x vs native MLX Attention, 20x prefill speedup on Mamba SSM
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



