Back to feed
arXiv cs.CL
arXiv cs.CL
7/23/2026
SLPO: Scaling Latent Reasoning via a Surrogate Policy

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Short summary

SLPO brings outcome-reward reinforcement learning to autoregressive latent reasoners, which carry intermediate computation as continuous vectors instead of decoded tokens. It introduces an empirical surrogate policy density for trajectory-level credit assignment and a correctness-supervised stopping head that refines into a variable-horizon policy. SLPO improves Pass@k under parallel sampling and allocates longer computation to harder instances with higher accuracy.

  • SLPO enables outcome-reward RL for latent reasoners, moving beyond imitation learning
  • Surrogate policy density enables trajectory-level credit assignment without per-step likelihoods
  • Correctness-supervised stopping head adapts computation length to problem difficulty

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more