Back to feed
arXiv cs.CL
arXiv cs.CL
7/22/2026
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

Short summary

Relay-Bench is a text-only benchmark that evaluates LLMs on composite multi-domain reasoning tasks within a single prompt. Problems combine two to thirteen subproblems spanning visual reasoning, coding, math, information extraction, and data analysis, with added context bloat. The leading model GPT-5.5 (xHigh) scores only 43.3%, indicating significant headroom for improvement.

  • Relay-Bench tests LLMs on composite multi-domain reasoning in single prompts
  • Problems combine 2-13 subproblems across coding, math, web search, data analysis, and more
  • Leading model GPT-5.5 (xHigh) scores 43.3%, showing the benchmark is far from saturated

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more