This paper introduces a new method, Surrogate Latent Policy Optimization (SLPO), to improve the performance of autoregressive latent reasoners in reinforcement learning by combining outcome-reward RL with latent reasoning. Practitioners might care because SLPO can help scale the performance of these models at test time, which is currently a major challenge.
Firehose
Filtered to Papers, tagged “outcome-reward RL” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News