Back to feed
Dev.to
Dev.to
7/26/2026
Claude Opus 5 Benchmarks: What the Numbers Actually Show

Claude Opus 5 Benchmarks: What the Numbers Actually Show

Short summary

Claude Opus 5 scores 79.2% on SWE-bench Pro, a 10-point jump from Opus 4.8 at unchanged per-token pricing, closing to within 1.1 points of Fable 5 at half the cost. The author argues Anthropic's launch page relied heavily on ratios rather than absolute scores, making independent verification difficult. SWE-bench Verified is now saturated above 90% for four models, making SWE-bench Pro the more meaningful discriminator for real coding ability.

  • Opus 5 gains 10 points on SWE-bench Pro over Opus 4.8 with no price increase
  • Anthropic published most gains as ratios, not absolute scores, complicating verification
  • SWE-bench Verified is saturated; SWE-bench Pro is now the meaningful benchmark for frontier models

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more