Back to feed
Dev.to
Dev.to
7/25/2026
The original title is "389 Tests Passed. NIST Still Caught the Bug."

The original title is "389 Tests Passed. NIST Still Caught the Bug."

Original: 389 Tests Passed. NIST Still Caught the Bug.

Short summary

The author tested an AI agent's calculator tool by mutating a multiplication operator to addition—389 of 390 tests still passed, exposing how conventional test suites miss subtle numerical bugs. Using NIST Statistical Reference Datasets and cargo-mutants, they demonstrate a rigorous calibration loop: claim, attack, fail, repair, mutate, replay. The key insight is that AI agents need a clear boundary between semantic interpretation and deterministic, contract-tested computation.

  • A single operator mutation passed 389/390 tests, revealing gaps in conventional test suites
  • NIST Statistical Reference Datasets provide certified values for validating computational correctness
  • The calibration loop—claim, attack, fail, repair, mutate, replay—ensures agent tools are inspectable and repeatable

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more