Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
Zhiyuan Zeng, Qinyuan Cheng, Zhangyue Yin, Yunhua Zhou, Xipeng Qiu
摘要
The advent of test-time scaling in large language models (LLMs), exemplified by Ope-nAI's o1 series, has advanced reasoning capabilities by scaling computational resource allocation during inference. While successors like QwQ, Deepseek-R1 (R1) and LIMO replicate these advancements, whether these models truly possess test-time scaling capabilities remains underexplored. This study found that longer CoTs of these o1-like models do not consistently enhance accuracy; in fact, correct solutions are often shorter than incorrect ones for the same questions. Further investigation shows this phenomenon is closely related to models' self-revision capabilities -longer CoTs contain more self-revisions, which often lead to performance degradation. We then compare sequential and parallel scaling strategies on QwQ, R1 and LIMO, finding that parallel scaling achieves better coverage and scalability. Based on these insights, we propose Shortest Majority Vote, a method that combines parallel scaling strategies with CoT length characteristics, significantly improving models' test-time scalability compared to conventional majority voting approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen 等NeurIPS 2025 · 被引用 177 次
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex TasksFali Wang, Hui Liu, Zhenwei Dai, Jingying Zeng 等NeurIPS 2025 · 被引用 20 次
- On Reasoning Strength Planning in Large Reasoning ModelsLeheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao 等NeurIPS 2025 · 被引用 17 次
- ATTS: Asynchronous Test-Time Scaling via Conformal PredictionJing Xiong, Qiujiang Chen, Fanghua Ye, Zhongwei Wan 等ICLR 2026 · 被引用 8 次
- Synthesizing Visual Concepts as Vision-Language ProgramsAntonia Wüst, Wolfgang Stammer, Hikaru Shindo, Lukas Helff 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
相关 Paper
- Slim-SC: Thought Pruning for Efficient Scaling with Self-ConsistencyColin Hong, Xu Guo, Anand Chaanan Singh, Esha Choukse 等EMNLP 2025
- Understanding the Role of Training Data in Test-Time ScalingAdel Javanmard, Baharan Mirzasoleiman, Vahab MirrokniICLR 2026 · 被引用 5 次
- Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning ModelsSoumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu 等NeurIPS 2025 · 被引用 43 次
- Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability TheoryYexiang Liu, Zekun Li, Zhi Fang, Nan Xu 等ACL 2025 · 被引用 12 次
- Let Me Think! A Long Chain of Thought Can Be Worth Exponentially Many Short OnesParsa Mirtaheri, Ezra Edelman, Samy Jelassi, Eran Malach 等NeurIPS 2025 · 被引用 12 次
