AR
arXiv CS.AI
7/24/2026

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
Short summary
InferenceBench is a benchmark where AI agents must deploy and optimize an OpenAI-compatible LLM inference server on a single H100 GPU within a two-hour budget. Across 15 frontier agent configurations, agents improve up to 8x over naive PyTorch but fall short of simple hyperparameter search at 11.5x. The key finding is that agents lack diverse configuration exploration rather than domain knowledge, revealing a bottleneck in open-ended AI engineering tasks.
- •Benchmark for evaluating AI agents on open-ended LLM inference optimization
- •Agents achieve up to 8x speedup but underperform hyperparameter search at 11.5x
- •Bottleneck is configuration diversity, not domain knowledge
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
