Back to feed
arXiv cs.LG
arXiv cs.LG
7/23/2026
Bayesian Wind Tunnels for Model Selection

Bayesian Wind Tunnels for Model Selection

Short summary

This paper introduces 'Bayesian wind tunnels' to test whether transformers can perform Bayesian model selection, not just filtering. A 2.8M-parameter transformer achieves near-optimal Bayesian model selection on involutions, but fails when the discriminative statistic requires arithmetic with opaque symbols—a boundary that persists even at 316M parameters. Frontier LLMs show qualitative Bayesian behavior but a large calibration gap (~55x).

  • Transformers can perform Bayesian model selection on relational tasks but fail when arithmetic is needed with opaque symbols
  • This failure boundary persists under 112x parameter scaling (2.8M to 316M)
  • Frontier LLMs show qualitative Bayesian behavior but a ~55x calibration gap

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more