Advancing LLM Reasoning Generalists with Preference Trees
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Boji Shan, Zeyuan Liu, Jia Deng, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu
Abstract
We introduce EURUS, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B and CodeLlama-70B, EU-RUS models achieve state-of-the-art results among open-source models on a diverse set of benchmarks covering mathematics, code generation, and logical reasoning problems. Notably, EURUS-70B beats GPT-3.5 Turbo in reasoning through a comprehensive benchmarking across 12 tests covering five tasks, and achieves a 33.3% pass@1 accuracy on LeetCode and 32.6% on TheoremQA, two challenging benchmarks, substantially outperforming existing open-source models by margins more than 13.3%. The strong performance of EU-RUS can be primarily attributed to ULTRAINTER-ACT, our newly-curated large-scale, high-quality alignment dataset specifically designed for complex reasoning tasks. ULTRAINTERACT can be used in both supervised fine-tuning and preference learning. For each instruction, it includes a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format, (2) multi-turn interaction trajectories with the environment and the critique, and (3) pairwise data to facilitate preference learning. UL-TRAINTERACT allows us to conduct an in-depth exploration of preference learning for reasoning tasks. Our investigation reveals that some wellestablished preference learning algorithms may be less suitable for reasoning tasks compared to their effectiveness in general conversations. In-* Equal contribution
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 744c325e-0c95-4636-bf1a-d32f59b42225Cited by top-tier papers97
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
- Recursive Introspection: Teaching Language Model Agents How to Self-ImproveYuxiao Qu, Tianjun Zhang, Naman Garg, Aviral KumarNeurIPS 2024 · 218 citations
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He et al.ICLR 2026 · 211 citations
- MAmmoTH2: Scaling Instructions from the WebXiang Yue, Tianyu Zheng, Ge Zhang, Wenhu ChenNeurIPS 2024 · 176 citations
- Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingYe Tian, Baolin Peng, Linfeng Song, Lifeng Jin et al.NeurIPS 2024 · 162 citations
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang et al.ICML 2024 · 436 citations
- What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction TuningWei Liu, Weihao Zeng, Keqing He, Yong Jiang et al.ICLR 2024 · 369 citations
Related papers
- Learning Planning-based Reasoning by Trajectories Collection and Process Reward SynthesizingFangkai Jiao, Chengwei Qin, Zhengyuan Liu, Nancy F. Chen et al.EMNLP 2024 · 1 citation
- Preference Optimization for Reasoning with Pseudo FeedbackFangkai Jiao, Geyang Guo, Xingxing Zhang, Nancy F. Chen et al.ICLR 2025
- DeepOR: A Deep Reasoning Foundation Model for Optimization ModelingZiyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan et al.AAAI 2026 · 1 citation
- Agent Instructs Large Language Models to be General Zero-Shot ReasonersNicholas Crispino, Kyle Montgomery, Fankun Zeng, Dawn Song et al.ICML 2024 · 41 citations
- TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference TreesWeibin Liao, Xu Chu, Yasha WangICLR 2025
