Dev.to
8/1/2026

Benchmarking AI Coding Agents on Real Pull Requests
Short summary
The authors built octobench, a benchmark of 25 real pull requests merged in 2026 across five languages, graded by each project's own held-out tests to avoid contamination and taste-grading issues in popular suites. Octomind with an open model solved 24/25 tasks, beating Claude Code with Opus, while the same model in another harness solved only 19 at double the cost. The benchmark is designed to be re-harvested as model training cutoffs advance, keeping it fresh and resistant to data contamination.
- •Built octobench: 25 real merged PRs across Python, PHP, Rust, C++, JS graded by project-held-out tests
- •Octomind with open model solved 24/25, beating Claude Code with Opus; same model in another harness solved 19 at 2x cost
- •Benchmark designed for re-harvesting as cutoffs advance to prevent contamination and staleness
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



