AR
arXiv CS.AI
7/24/2026

The original title is "JAXBench: Benchmarking Autonomous TPU Kernel Optimization"
Original: JAXBench: Benchmarking Autonomous TPU Kernel Optimization
Short summary
JAXBench is a TPU-native benchmark suite of 50 JAX workloads for evaluating AI-generated kernel optimization on Google Cloud TPUs, derived from production ML operators in Llama-3.1, DeepSeek-V3, and Mixtral. With Gemini 3 Flash, conditioning on curated TPU documentation raised correctness from 5.8% to 37.3% and solved 48 of 50 benchmarks at 1.28x geomean speedup. Autocomp's beam-search pipeline achieved 1.36x geomean over XLA, recovering most of the 2.08x expert upper bound.
- •50-workload TPU benchmark suite for AI-generated kernel optimization using JAX and Pallas
- •Target-specific documentation matters more than model scale: correctness jumps from 5.8% to 37.3% with curated TPU docs
- •Autocomp beam-search achieves 1.36x geomean speedup over XLA, approaching expert hand-tuned baselines
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

