Back to feed
Dev.to
Dev.to
7/21/2026
Kimi K3 vs Claude Fable 5 and Opus 4.8: a benchmark you can run yourself

Kimi K3 vs Claude Fable 5 and Opus 4.8: a benchmark you can run yourself

Short summary

An open-source benchmark compares Kimi K3, Claude Fable 5, and Claude Opus 4.8 on a double-entry ledger task with hidden money-safety tests that the models never see. All three models fixed the known bugs and added features correctly across nine runs, but differentiated on long-term code quality choices like encapsulation enforcement and defensive testing. The benchmark focuses on code maintainability rather than simple pass/fail results.

  • Open-source benchmark tests Kimi K3, Claude Fable 5, and Claude Opus 4.8 on a ledger task with hidden money-safety tests
  • All three models fixed known bugs and added features across nine runs — differentiation was in code quality and defensive practices
  • Claude Opus 4.8 was most consistent, adding self-tests for encapsulation invariants; benchmark is runnable at github.com/dsplce-co/kimi-vs-fable-vs-opus

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more