Back to feed
AR
arXiv CS.AI
7/23/2026
Information Discernment in Large Language Models

Information Discernment in Large Language Models

Short summary

Researchers introduce Learn2Discern (L2D), a benchmark testing whether LLMs appropriately weigh external information by source reliability and truth alignment. Across 13 models and 670K trials, models perform near chance on both source and truth discernment, relying on popularity over reliability. Newer/larger models improve truth discernment but not source discernment, though simple inference-time interventions help both.

  • L2D benchmark formalizes information discernment via three normative axioms validated by a 299-person user study
  • Models rely on source popularity 2x more than reliability and update equally whether claims improve or worsen accuracy
  • Inference-time interventions improve both discernment forms; dataset and survey released as public testbed

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more