Back to feed
AR
arXiv CS.AI
7/13/2026
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Short summary

Long-Horizon-Terminal-Bench introduces 46 long-horizon terminal tasks across nine categories with dense reward-based grading that captures intermediate progress, not just final outcomes. Tasks require hundreds of episodes and minutes to hours of execution, with frontier models consuming ~9.9M tokens per task on average. Even the strongest model achieves only 15.2% pass@1 at 0.95 partial-reward threshold, revealing significant headroom for improvement in long-horizon agent capabilities.

  • 46 long-horizon terminal tasks with dense intermediate rewards and partial credit
  • Frontier models average 9.9M tokens/task, 231 episodes, 85.3 minutes per run
  • Best model achieves only 15.2% pass@1 at 0.95 threshold, showing major improvement headroom

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more