Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models
Jonathan Williams, Olga Russakovsky, Esin Tureci
Abstract
Looped Language Models (LoopLMs) perform multi-step latent reasoning prior to token generation and outperform conventional LLMs on reasoning benchmarks at smaller parameter budgets. However, attempts to further improve LoopLM reasoning with reinforcement learning have failed—standard objectives such as Group Relative Policy Optimization (GRPO) only assign credit to the final latent state, creating a fundamental mismatch with the model's internal computation. To resolve this, we introduce RLTT (Reward Latent Thought Trajectories), a reinforcement learning framework which distributes reward across the full latent reasoning trajectory. RLTT provides dense, trajectory-level credit assignment without relying on external verifiers and can directly replace GRPO with negligible overhead. Across extensive experiments with Ouro-1.4B/2.6B-Thinking under identical training and inference conditions, RLTT yields statistically significant improvements over GRPO on challenging mathematical reasoning benchmarks, improving mean accuracy over MATH-500, AIME24/26, and BeyondAIME by +5.8% on the 1.4B scale, and +10.9% on the 2.6B scale. Despite being trained exclusively on mathematics, RLTT also transfers effectively to non-mathematical reasoning benchmarks, demonstrating the effectiveness of trajectory-level credit assignment for reinforcement learning in LoopLMs. Code is available at https://github.com/jonwill8/RLTT.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fb08d30-781d-4aea-8ac2-3ef66dffc858Builds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer et al.NeurIPS 2025 · 431 citations
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen et al.NeurIPS 2025 · 177 citations
- Looped Transformers as Programmable ComputersAngeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee et al.ICML 2023 · 175 citations
- Looped Transformers are Better at Learning Learning AlgorithmsLiu Yang, Kangwook Lee, Robert D. Nowak, Dimitris PapailiopoulosICLR 2024 · 82 citations
Related papers
- Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy OptimizationJunzhe Wang, Zhiheng Xi, Yajie Yang, Hao Luo et al.ACL 2026 · 4 citations
- Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM ReasoningZhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen et al.ICML 2026 · 24 citations
- Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement LearningWeijie Li, Jin Wang, Liang-Chih Yu, Xuejie ZhangAAAI 2026 · 1 citation
- GRPO is Secretly a Process Reward ModelMichael Sullivan, Alexander KollerICML 2026 · 8 citations
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin et al.ICML 2026 · 53 citations
