Dev.to
7/21/2026

Kimi K3 vs Claude Fable 5 and Opus 4.8: a benchmark you can run yourself
Short summary
An open-source benchmark compares Kimi K3, Claude Fable 5, and Claude Opus 4.8 on a double-entry ledger task with hidden money-safety tests that the models never see. All three models fixed the known bugs and added features correctly across nine runs, but differentiated on long-term code quality choices like encapsulation enforcement and defensive testing. The benchmark focuses on code maintainability rather than simple pass/fail results.
- •Open-source benchmark tests Kimi K3, Claude Fable 5, and Claude Opus 4.8 on a ledger task with hidden money-safety tests
- •All three models fixed known bugs and added features across nine runs — differentiation was in code quality and defensive practices
- •Claude Opus 4.8 was most consistent, adding self-tests for encapsulation invariants; benchmark is runnable at github.com/dsplce-co/kimi-vs-fable-vs-opus
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



