Parallel Test-Time Scaling for Latent Reasoning Models
Runyang You, Yongqi Li, Meng Liu, Wenjie Wang, Liqiang Nie, Wenjie Li
Abstract
Parallel test-time scaling (TTS) is a pivotal approach for enhancing large language models (LLMs), typically by sampling multiple token-based chains-of-thought in parallel and aggregating outcomes through voting or search. Recent advances in latent reasoning, where intermediate reasoning unfolds in continuous vector spaces, offer a more efficient alternative to explicit Chain-of-Thought, yet whether such latent models can similarly benefit from parallel TTS remains open, mainly due to the absence of sampling mechanisms in continuous space, and the lack of probabilistic signals for advanced trajectory aggregation. This work enables parallel TTS for latent reasoning models by addressing the above issues. For sampling, we introduce two uncertainty-inspired stochastic strategies: Monte Carlo Dropout and Additive Gaussian Noise. For aggregation, we design a Latent Reward Model (LatentRM) trained with step-wise contrastive objective to score and guide latent reasoning. Extensive experiments and visualization analyses show that both sampling strategies scale effectively with compute and exhibit distinct exploration dynamics, while LatentRM enables effective trajectory selection. Together, our explorations open a new direction for scalable inference in continuous spaces. Code and checkpoints released at https://github.com/ModalityDance/LatentTTS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4053f5d4-a4bb-4dee-82a4-3bd96a117aa9Cited by top-tier papers5
- The Emergence of Abstract Thought in Large Language Models Beyond Any LanguageYuxin Chen, Yiran Zhao, Yang Zhang, An Zhang et al.NeurIPS 2025 · 27 citations
- Equilibrium Reasoners: Learning Attractors Enables Scalable ReasoningBenhao Huang, Zhengyang Geng, Zico KolterICML 2026 · 5 citations
- Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference OptimizationXinyu Qiu, Heng Jia, Zhengwen Zeng, Shuheng Shen et al.CVPR 2026 · 4 citations
- Verifiable Reasoning for LLM-based Generative RecommendationXinyu Lin, Hanqing Zeng, Hanchao Yu, Yinglong Xia et al.SIGIR 2026 · 1 citation
- What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic AnalysisXinghao Chen, Chak Tou Leong, Wenjin Guo, Jian Wang et al.ICML 2026 · 1 citation
Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft ThinkingJunhong Wu, Jinliang Lu, Zixuan Ren, Gangqiang Hu et al.ICLR 2026 · 35 citations
Related papers
- A Formal Comparison Between Chain of Thought and Latent ThoughtKevin Xu, Issei SatoICML 2026 · 12 citations
- UniT: Unified Multimodal Chain-of-Thought Test-time ScalingLeon Liangyu Chen, Haoyu Ma, Zhipeng Fan, Ziqi Huang et al.CVPR 2026 · 7 citations
- Vision-aligned Latent Reasoning for Multi-modal Large Language ModelByungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho et al.ICML 2026 · 7 citations
- Slim-SC: Thought Pruning for Efficient Scaling with Self-ConsistencyColin Hong, Xu Guo, Anand Chaanan Singh, Esha Choukse et al.EMNLP 2025
- Enhancing Language Model Reasoning with Structured Multi-Level ModelingSiheng Xiong, Ali Payani, Faramarz FekriICLR 2026
