Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
Haozhe Wang, Qixin Xu, Che Liu, Junhong Wu, Fangzhen Lin, Wenhu Chen
摘要
Reinforcement Learning (RL) has proven highly effective at enhancing the complex reasoning abilities of Large Language Models (LLMs), yet underlying mechanisms driving this success remain largely opaque. Our analysis reveals that puzzling phenomena like "aha moments", "length-scaling" and entropy dynamics are not disparate occurrences but hallmarks of an emergent reasoning hierarchy, akin to the separation of high-level strategic planning from low-level procedural execution in human cognition. We uncover a compelling two-phase dynamic: initially, a model is constrained by procedural correctness and must improve its low-level skills. The learning bottleneck then decisively shifts, with performance gains being driven by the exploration and mastery of high-level strategic planning. This insight exposes a core inefficiency in prevailing RL algorithms like GRPO, which apply optimization pressure agnostically and dilute the learning signal across all tokens. To address this, we propose Hierarchy-Aware Credit Assignment (HICRA), an algorithm that concentrates optimization efforts on high-impact planning tokens. low-level procedural executions. (Right) Hierarchical reasoning emerges during RL training via a two-phase dynamic. Phase ① consolidates low-level skills, marked by a token-entropy drop in execution tokens. The learning frontier then shifts to Phase ②, where the model explores and masters high-level planning, marked by increased semantic diversity, sustained reasoning enhancement and length scaling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu 等NeurIPS 2025 · 被引用 356 次
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning MemorySiru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen 等ICLR 2026 · 被引用 244 次
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationQingyang Zhang, Haitao Wu, Changqing Zhang, Peilin Zhao 等NeurIPS 2025 · 被引用 134 次
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensRuichuan An, Sihan Yang, Renrui Zhang, Zijun Shen 等NeurIPS 2025 · 被引用 61 次
- Reverse-Engineered Reasoning for Open-Ended GenerationHaozhe Wang, Haoran Que, Qixin Xu, Minghao Liu 等ICLR 2026 · 被引用 38 次
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu 等NeurIPS 2025 · 被引用 356 次
相关 Paper
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin 等ICML 2026 · 被引用 53 次
- Learning to Reason over Continuous Tokens with Reinforcement LearningYiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong 等ICLR 2026 · 被引用 1 次
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han 等ICLR 2026 · 被引用 47 次
- Attention Illuminates LLM Reasoning: The Uncovered Preplan-and-Anchor Rhythm Enables Fine-Grained Policy OptimizationYang Li, Zhichen Dong, Yuhan Sun, Weixun Wang 等ICML 2026 · 被引用 25 次
- Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM ReasoningZhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen 等ICML 2026 · 被引用 24 次
