arXiv cs.CL
7/24/2026

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces
Short summary
Researchers investigate how reasoning-enabled LLMs judge literary quality by extracting implicit theories from reasoning traces. Using a 30-text benchmark across six quality tiers, DeepSeek achieves 79.3% tier-classification accuracy, valuing intentionality, craft, depth, and distinctive voice. Degradation experiments show LLM judgments are holistic and more sensitive to structural features (voice, structure) than lexical ones (vocabulary), with implications for automated writing feedback.
- •DeepSeek achieves 79.3% accuracy classifying texts across six literary quality tiers
- •LLMs value intentionality, craft, depth, and voice over surface-level correctness
- •Structural degradation (voice, structure) hurts quality scores far more than vocabulary simplification
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
