Back to feed
arXiv cs.CL
arXiv cs.CL
7/23/2026
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

Short summary

A new reference-free framework audits LLM reasoning by decomposing traces into segments, labeling premise-target relations via NLI, and organizing them into a hypergraph for deterministic backward search. Evaluated on mathematical reasoning (Hard2Verify) and a new physician-annotated clinical benchmark (UroReason), the NLI-hypergraph method outperforms LLM-as-judge baselines, which often over-accept fluent but weakly grounded responses.

  • NLI-hypergraph framework for reference-free auditing of LLM reasoning traces
  • Introduces UroReason, a physician-annotated benchmark for clinical LLM reasoning
  • Outperforms LLM-as-judge baselines that over-accept fluent but weakly grounded outputs

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more