Back to feed
arXiv cs.CL
arXiv cs.CL
7/23/2026
When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

Short summary

This study examines how supervised fine-tuning (SFT) reduces behavioral diversity in LLM sequential decision-making, using deterministic board games as a controlled testbed. The authors find that reasoning-mode generation suppresses action diversity without uniformly improving accuracy, and standard SFT causes premature diversity collapse beyond what the accuracy-diversity tradeoff requires. Action augmentation—training on all optimal actions per state rather than a single demonstration—partially mitigates this effect.

  • SFT causes premature diversity collapse in LLM game play beyond accuracy-diversity tradeoff needs
  • Reasoning-mode generation suppresses action diversity without uniformly improving accuracy
  • Action augmentation (training on all optimal actions) partially mitigates diversity loss

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more