arXiv cs.CL
7/23/2026

SLPO: Scaling Latent Reasoning via a Surrogate Policy
Short summary
SLPO brings outcome-reward reinforcement learning to autoregressive latent reasoners, which carry intermediate computation as continuous vectors instead of decoded tokens. It introduces an empirical surrogate policy density for trajectory-level credit assignment and a correctness-supervised stopping head that refines into a variable-horizon policy. SLPO improves Pass@k under parallel sampling and allocates longer computation to harder instances with higher accuracy.
- •SLPO enables outcome-reward RL for latent reasoners, moving beyond imitation learning
- •Surrogate policy density enables trajectory-level credit assignment without per-step likelihoods
- •Correctness-supervised stopping head adapts computation length to problem difficulty
Generated with AI, which can make mistakes.
Is this a good recommendation for you?