Reasoning with Exploration: An Entropy Perspective
Daixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai, Xin Zhao, Zhenliang Zhang, Furu Wei
摘要
Balancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing language model (LM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting deeper and longer reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LM reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper89
- R-Zero: Self-Evolving Reasoning LLM from Zero DataChengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang 等ICLR 2026 · 被引用 220 次
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 等ICLR 2026 · 被引用 146 次
- Geometric-Mean Policy OptimizationYuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen 等ICLR 2026 · 被引用 104 次
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest QuestionsLu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang 等ICLR 2026 · 被引用 103 次
- Entropy-Aware On-Policy Distillation of Language ModelsWoogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei 等ICML 2026 · 被引用 91 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren 等NeurIPS 2025 · 被引用 314 次
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 被引用 304 次
相关 Paper
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo 等ACL 2026 · 被引用 42 次
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin 等ICML 2026 · 被引用 53 次
- Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's NatureZheng Liu, Mengjie Liu, Siwei Wen, Mengzhang Cai 等ACL 2026 · 被引用 9 次
- CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement LearningZhenpeng Su, Leiyu Pan, Minxuan Lv, Yuntao Li 等ACL 2026 · 被引用 21 次
- Emergent Hierarchical Reasoning in LLMs through Reinforcement LearningHaozhe Wang, Qixin Xu, Che Liu, Junhong Wu 等ICLR 2026 · 被引用 44 次
