Back to feed
arXiv cs.LG
arXiv cs.LG
7/24/2026
Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts

Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts

Short summary

This paper dissects adaptive depth in looped Transformers, showing that learned halting gates entangle exit selection with trajectory formation, causing poor adaptive-compute performance. Through controlled synthetic tasks and large-scale Ouro-1.4B/2.6B checkpoints, the authors find that fixed-prior depth supervision produces difficulty-aware trajectories and simple post-hoc confidence readouts often match or outperform learned gates. The failure stems from joint gate training's effect on trajectories, not limited gate expressivity, reframing adaptive depth as a joint trajectory-formation and exit-readout problem.

  • Learned halting gates entangle exit selection with trajectory supervision in looped Transformers
  • Fixed-prior depth supervision plus simple post-hoc confidence readouts can match or beat learned gates
  • Failure localizes to trajectory induced by joint gate training, not gate expressivity limitations

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more