Dev.to
7/25/2026

The original title is "389 Tests Passed. NIST Still Caught the Bug."
Original: 389 Tests Passed. NIST Still Caught the Bug.
Short summary
The author tested an AI agent's calculator tool by mutating a multiplication operator to addition—389 of 390 tests still passed, exposing how conventional test suites miss subtle numerical bugs. Using NIST Statistical Reference Datasets and cargo-mutants, they demonstrate a rigorous calibration loop: claim, attack, fail, repair, mutate, replay. The key insight is that AI agents need a clear boundary between semantic interpretation and deterministic, contract-tested computation.
- •A single operator mutation passed 389/390 tests, revealing gaps in conventional test suites
- •NIST Statistical Reference Datasets provide certified values for validating computational correctness
- •The calibration loop—claim, attack, fail, repair, mutate, replay—ensures agent tools are inspectable and repeatable
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



