Back to feed
AR
arXiv CS.AI
7/22/2026
Calibrated Selective Fact-Checking via Evidence Chain Evaluation

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

Short summary

Evidence Chain Evaluation (ECE) is a selective fact-checking framework that lets LLM-based verification agents abstain from binary true/false verdicts when evidence is weak or inconsistent. On ECE-Bench it achieves 91.6% standard accuracy and 97.8% selective accuracy on answered claims, deferring only 6 of 95 cases concentrated in low-reliability evidence settings. While it doesn't beat the strongest retrieval baseline on aggregate calibration metrics, it demonstrates that abstention functions as a safety mechanism for epistemically weak evidence.

  • ECE framework allows verification agents to return 'uncertain' instead of forced true/false verdicts
  • Achieves 97.8% selective accuracy on answered claims with 93.7% coverage
  • Deferred cases cluster in low-reliability evidence settings, validating abstention as a safety mechanism

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more