Back to feed
Dev.to
Dev.to
8/1/2026
Benchmarking AI Coding Agents on Real Pull Requests

Benchmarking AI Coding Agents on Real Pull Requests

Short summary

The authors built octobench, a benchmark of 25 real pull requests merged in 2026 across five languages, graded by each project's own held-out tests to avoid contamination and taste-grading issues in popular suites. Octomind with an open model solved 24/25 tasks, beating Claude Code with Opus, while the same model in another harness solved only 19 at double the cost. The benchmark is designed to be re-harvested as model training cutoffs advance, keeping it fresh and resistant to data contamination.

  • Built octobench: 25 real merged PRs across Python, PHP, Rust, C++, JS graded by project-held-out tests
  • Octomind with open model solved 24/25, beating Claude Code with Opus; same model in another harness solved 19 at 2x cost
  • Benchmark designed for re-harvesting as cutoffs advance to prevent contamination and staleness

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more