Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
Yulei Qin, Gang Li, Zongyi Li, Zihan Xu, Yuchen Shi, Zhekai Lin, Xiao Cui, Ke Li, Xing Sun
摘要
Existing large language models (LLMs) face challenges of following complex instructions, especially when multiple constraints are present and organized in paralleling, chaining, and branching structures. One intuitive solution, namely chain-of-thought (CoT), is expected to universally improve capabilities of LLMs. However, we find that the vanilla CoT exerts a negative impact on performance due to its superficial reasoning pattern of simply paraphrasing the instructions. It fails to peel back the compositions of constraints for identifying their relationship across hierarchies of types and dimensions. To this end, we propose RAIF, a systematic method to boost LLMs in dealing with complex instructions via incentivizing reasoning for test-time compute scaling. First, we stem from the decomposition of complex instructions under existing taxonomies and propose a reproducible data acquisition method. Second, we exploit reinforcement learning (RL) with verifiable rule-centric reward signals to cultivate reasoning specifically for instruction following. We address the shallow, non-essential nature of reasoning under complex instructions via sample-wise contrast for superior CoT enforcement. We also exploit behavior cloning of experts to facilitate steady distribution shift from fast-thinking LLMs to skillful reasoners. Extensive evaluations on seven comprehensive benchmarks confirm the validity of the proposed method, where a 1.5B LLM achieves 11.74% gains with performance comparable to a 8B LLM. Evaluation on OOD constraints also confirms the generalizability of our RAIF. Codes and data are available at https://github.com/yuleiqin/RAIF. Keywords: reinforcement learning with verifiable rewards (RLVR), instruction following, complex instructions
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningYulei Qin, Xiaoyu Tan, Zhengbao He, Gang Li 等ICLR 2026 · 被引用 9 次
- RAIN-Merging: A Gradient-Free Method to Enhance Instruction Following in Large Reasoning Models with Preserved Thinking FormatZhehao Huang, Yuhang Liu, Baijiong Lin, Yixin Lou 等ICLR 2026 · 被引用 7 次
- Instructions are all you need: Self-supervised Reinforcement Learning for Instruction FollowingQingyu Ren, Qianyu He, Powei Chang, Jie Zeng 等ACL 2026 · 被引用 6 次
- Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction FollowingKongcheng Zhang, QI YAO, Shunyu Liu, Wenjian Zhang 等ICML 2026 · 被引用 4 次
- ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction FollowingYuancheng Yang, Lin Yang, Xu Wang, Chao Tong 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper36
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack EfficientlyStanley Wei, Juno KimICML 2026
- VerIF: Verification Engineering for Reinforcement Learning in Instruction FollowingHao Peng, Yunjia Qi, Xiaozhi Wang, Bin Xu 等EMNLP 2025 · 被引用 24 次
- Demystifying Long Chain-of-Thought ReasoningShiming Yang, Yuxuan Tong, Xinyao Niu, Graham Neubig 等ICML 2025
- Rectifying LLM Thought from Lens of OptimizationJunnan Liu, Hongwei Liu, Songyang Zhang, Kai ChenICLR 2026 · 被引用 3 次
- When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMsXiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen 等NeurIPS 2025 · 被引用 63 次
