AR
arXiv CS.AI
7/23/2026

Information Discernment in Large Language Models
Short summary
Researchers introduce Learn2Discern (L2D), a benchmark testing whether LLMs appropriately weigh external information by source reliability and truth alignment. Across 13 models and 670K trials, models perform near chance on both source and truth discernment, relying on popularity over reliability. Newer/larger models improve truth discernment but not source discernment, though simple inference-time interventions help both.
- •L2D benchmark formalizes information discernment via three normative axioms validated by a 299-person user study
- •Models rely on source popularity 2x more than reliability and update equally whether claims improve or worsen accuracy
- •Inference-time interventions improve both discernment forms; dataset and survey released as public testbed
Generated with AI, which can make mistakes.
Is this a good recommendation for you?