arXiv cs.CL
7/22/2026

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
Short summary
Relay-Bench is a text-only benchmark that evaluates LLMs on composite multi-domain reasoning tasks within a single prompt. Problems combine two to thirteen subproblems spanning visual reasoning, coding, math, information extraction, and data analysis, with added context bloat. The leading model GPT-5.5 (xHigh) scores only 43.3%, indicating significant headroom for improvement.
- •Relay-Bench tests LLMs on composite multi-domain reasoning in single prompts
- •Problems combine 2-13 subproblems across coding, math, web search, data analysis, and more
- •Leading model GPT-5.5 (xHigh) scores 43.3%, showing the benchmark is far from saturated
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
