arXiv cs.LG
7/24/2026

Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts
Short summary
This paper dissects adaptive depth in looped Transformers, showing that learned halting gates entangle exit selection with trajectory formation, causing poor adaptive-compute performance. Through controlled synthetic tasks and large-scale Ouro-1.4B/2.6B checkpoints, the authors find that fixed-prior depth supervision produces difficulty-aware trajectories and simple post-hoc confidence readouts often match or outperform learned gates. The failure stems from joint gate training's effect on trajectories, not limited gate expressivity, reframing adaptive depth as a joint trajectory-formation and exit-readout problem.
- •Learned halting gates entangle exit selection with trajectory supervision in looped Transformers
- •Fixed-prior depth supervision plus simple post-hoc confidence readouts can match or beat learned gates
- •Failure localizes to trajectory induced by joint gate training, not gate expressivity limitations
Generated with AI, which can make mistakes.
Is this a good recommendation for you?