PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
Ruishuo Chen, Yu Chen, Zhuoran Li, Longbo Huang
摘要
Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theoretical optimization target and are prone to degenerative biases. In this work, we introduce PowerFlow, a principled framework that reformulates unsupervised fine-tuning as a distribution matching problem. By casting GFlowNet as an amortized variational sampler for unnormalized densities, we propose a length-aware Trajectory-Balance objective that explicitly neutralizes the structural length biases inherent in autoregressive generation. By targeting -power distributions, PowerFlow enables the directional elicitation of the dual nature of LLMs: sharpening the distribution () to intensify logical reasoning, or flattening it () to unlock expressive creativity. Extensive experiments demonstrate that PowerFlow consistently outperforms existing RLIF methods, matching or even exceeding supervised GRPO. Furthermore, by mitigating over-sharpening in aligned models, our approach achieves simultaneous gains in diversity and quality, shifting the Pareto frontier in creative tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
相关 Paper
- FlowRL: Matching Reward Distributions for LLM ReasoningXuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li 等ICLR 2026 · 被引用 41 次
- Amortizing intractable inference in large language modelsEdward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar 等ICLR 2024 · 被引用 91 次
- Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMsYiming Huang, Zhenbo Shi, Xin-Cheng Wen, Jichuan Zeng 等ACL 2026
- Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal ExamplesFangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao 等ICML 2025
- Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive FeedbackQian Dong, Yiding Liu, Qingyao Ai, Zhijing Wu 等SIGIR 2024 · 被引用 9 次
