Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
Yiming Huang, Zhenbo Shi, Xin-Cheng Wen, Jichuan Zeng, Cuiyun Gao, Peiyi Han, Chuanyi Liu
摘要
Unsupervised reinforcement learning (RL) has emerged as a promising paradigm for enabling self-improvement in large language models (LLMs). However, existing unsupervised RL-based methods often lack the capacity to adapt to the model's evolving reasoning capabilities during training. Therefore, these methods can misdirect policy optimization in the absence of ground-truth supervision. To address this issue, we introduce FREIA, a novel RL-based algorithm built on two key innovations: (1) Free Energy-Driven Reward (FER) adapts rewards to balance consensus and exploration based on the Free Energy Principle. (2) Adaptive Advantage Shaping (AAS) adaptively adjusts learning signals based on the statistical characteristics of sampled rewards. Empirical evaluations on nine datasets across three reasoning tasks showcase that FREIA outperforms other unsupervised RL-based baselines. Notably, in mathematical reasoning tasks, FREIA surpasses other methods by an average of 0.5 to 3.5 points in Pass@1 using the DeepSeek-R1-Distill-Qwen-1.5B model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye 等ICLR 2026 · 被引用 279 次
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine 等ICLR 2026 · 被引用 218 次
- The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningShivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han 等NeurIPS 2025 · 被引用 185 次
- Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic ReasoningPan Lu, Ran Gong, Shibiao Jiang, Liang Qiu 等ACL 2021
相关 Paper
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationQingyang Zhang, Haitao Wu, Changqing Zhang, Peilin Zhao 等NeurIPS 2025 · 被引用 134 次
- Reinforcing General Reasoning Without VerifiersXiangxin Zhou, Zichen Liu, Anya Sims, Haonan Wang 等ICLR 2026 · 被引用 75 次
- Prototype Entropy Alignment: Reinforcing Structured Uncertainty in LLM ReasoningZhengyuan Pan, Yanhao Chen, Zhongquan Jian, Wanru Zhao 等AAAI 2026
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM ReasoningKongcheng Zhang, Qi Yao, Shunyu Liu, Yingjie Wang 等NeurIPS 2025 · 被引用 45 次
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu 等NeurIPS 2025 · 被引用 249 次
