Mode-conditioning unlocks superior test-time compute scaling
Chen Henry Wu, Sachin Goyal, Aditi Raghunathan
摘要
Parallel sampling is essential to test-time scaling and reinforcement learning (RL), but its effectiveness is sharply limited by diversity collapse, where models concentrate on a few modes and repeated samples produce the same mistakes. We propose the mode-conditioning (ModC) framework, which explicitly allocates sampling compute across reasoning modes using either specialist models or mode-specific prefixes. With predefined mode labels, ModC consistently improves test-time scaling (Pass@k) across controlled graph-search tasks and math reasoning benchmarks, spanning model families and sizes from 0.5B to 7B. On OpenThoughts, fine-tuning Qwen2.5-7B with ModC achieves an 4× efficiency gain over standard training while also improving the maximum attainable Pass@k. We further show that gradient clustering enables ModC without predefined mode labels, yielding up to 10% gains on datasets such as NuminaMath. Finally, we show that ModC improves Pass@k after RL training and can further boost the Pass@k gains of diversity-inducing RL methods. These results demonstrate that standard training underutilizes the diversity in data, and that ModC provides a simple, effective remedy for unlocking the full benefits of diversity in parallel sampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
- OpenThoughts: Data Recipes for Reasoning ModelsEtash Kumar Guha, Ryan Marten, Sedrick Keh, Negin Raoof 等ICLR 2026 · 被引用 235 次
- Art or Artifice? Large Language Models and the False Promise of CreativityTuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan 等CHI 2024 · 被引用 122 次
相关 Paper
- Learning Diverse Responses with Prefix-Conditioned Supervised Fine-TuningZhiyuan Fan, Guanqiao Chen, Yanyi Huang, Mingkuan Zhao 等ACL 2026
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang 等ACL 2026 · 被引用 17 次
- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo MethodsIsha Puri, Shivchander Sudalairaj, Guangxuan Xu, Abhishek Bhandwaldar 等NeurIPS 2025 · 被引用 8 次
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated ReasoningJingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang 等ACL 2026 · 被引用 15 次
- Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning ModelsSoumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu 等NeurIPS 2025 · 被引用 43 次
