AR
arXiv CS.AI
7/13/2026

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Short summary
Long-Horizon-Terminal-Bench introduces 46 long-horizon terminal tasks across nine categories with dense reward-based grading that captures intermediate progress, not just final outcomes. Tasks require hundreds of episodes and minutes to hours of execution, with frontier models consuming ~9.9M tokens per task on average. Even the strongest model achieves only 15.2% pass@1 at 0.95 partial-reward threshold, revealing significant headroom for improvement in long-horizon agent capabilities.
- •46 long-horizon terminal tasks with dense intermediate rewards and partial credit
- •Frontier models average 9.9M tokens/task, 231 episodes, 85.3 minutes per run
- •Best model achieves only 15.2% pass@1 at 0.95 threshold, showing major improvement headroom
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
