Back to feed
arXiv cs.LG
arXiv cs.LG
7/24/2026
PhantomFill: When the Form Demands an Answer, Language Models Invent One

PhantomFill: When the Form Demands an Answer, Language Models Invent One

Short summary

PhantomFill demonstrates that requiring LLMs to fill structured form fields (JSON, enums, arrays) causes systematic hallucination even when inputs lack the necessary information. Across 13 models, required fields drove fabrication to 100% in 10 of them; GPT-5.5 answered honestly 98% of the time in free text but fabricated answers 40 out of 40 times when given a required JSON field. The benchmark reports Coerced Fabrication Rate and Escape Utilization Rate, and shows a one-line schema fix can mitigate the issue.

  • Required JSON fields cause 100% fabrication in 10 of 13 tested models when answers don't exist in the input
  • Explicit 'insufficient evidence' options only help frontier models; all 9 open-weight models ignore them
  • PhantomFill benchmark provides deterministic scoring with Coerced Fabrication Rate and Escape Utilization Rate metrics

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more