Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
Huifeng Yin, Yu Zhao, Minghao Wu, Xuanfan Ni, Bo Zeng, Hao Wang, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang
摘要
Large Reasoning Models (LRMs) such as Ope-nAI o1 and DeepSeek-R1 have shown remarkable reasoning capabilities by scaling testtime compute and generating long Chain-of-Thought (CoT). Distillation post-training on LRMs-generated data is a straightforward yet effective method to enhance the reasoning abilities of smaller models, but faces a critical bottleneck: we found that distilled long CoT data poses learning difficulty for small models and leads to the inheritance of biases (i.e. over-thinking) when using Supervised Fine-tuning (SFT) and Reinforcement Learning (RL) methods. To alleviate this bottleneck, we propose constructing tree-based CoT data from scratch via Monte Carlo Tree Search (MCTS). We then exploit a set of CoTaware approaches, including Thoughts Length Balance, Fine-grained DPO, and Joint Posttraining Objective, to enhance SFT and RL on the constructed data. We conduct evaluation on various benchmarks such as math (GSM8K, MATH, AIME). instruction-following (Multi-IF) and planning (Blocksworld), results demonstrate our approaches substantially improve the reasoning performance of distilled models compared to standard distilled models via reducing the hallucinations in long-time thinking. The project homepage is https://github.com/ AIDC-AI/Marco-o1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- QFFT, Question-Free Fine-Tuning for Adaptive ReasoningWanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin 等NeurIPS 2025 · 被引用 26 次
- Long-Chain Reasoning Distillation via Adaptive Prefix AlignmentZhenghao Liu, Zhuoyang Wu, Xinze Li, Yukun Yan 等ACL 2026 · 被引用 4 次
它引用的顶会 Paper5
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
- Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingYe Tian, Baolin Peng, Linfeng Song, Lifeng Jin 等NeurIPS 2024 · 被引用 162 次
- Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL DivergenceJunru Lu, Jiazheng Li, Siyu An, Meng Zhao 等EMNLP 2024 · 被引用 2 次
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsPeiyi Wang, Lei Li, Zhihong Shao, Runxin Xu 等ACL 2024
相关 Paper
- Enhancing Language Model Reasoning with Structured Multi-Level ModelingSiheng Xiong, Ali Payani, Faramarz FekriICLR 2026
- The Markovian Thinker: Architecture-Agnostic Linear Scaling of ReasoningMilad Aghajohari, Kamran Chitsaz, Amirhossein Kazemnejad, Sarath Chandar 等ICLR 2026
- Teaching Language Models to Reason with ToolsChengpeng Li, Zhengyang Tang, Ziniu Li, Mingfeng Xue 等NeurIPS 2025 · 被引用 8 次
- R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?Yi Lu, Jianing Wang, Linsen Guo, Wei He 等ICLR 2026 · 被引用 9 次
- VeriThinker: Learning to Verify Makes Reasoning Model EfficientZigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu 等NeurIPS 2025 · 被引用 30 次
