EXPO: Stable Reinforcement Learning with Expressive Policies
Perry Dong, Qiyang Li, Dorsa Sadigh, Chelsea Finn
摘要
We study the problem of training and fine-tuning expressive policies with online reinforcement learning (RL) given an offline dataset. Training expressive policy classes with online RL present a unique challenge of stable value maximization. Unlike simpler Gaussian policies commonly used in online RL, expressive policies like diffusion and flow-matching policies are parameterized by a long denoising chain, which hinders stable gradient propagation from actions to policy parameters when optimizing against some value function. Our key insight is that we can address stable value maximization by avoiding direct optimization over value with the expressive policy and instead construct an on-the-fly RL policy to maximize Q-value. We propose Expressive Policy Optimization (EXPO), a sample-efficient online RL algorithm that utilizes an on-the-fly policy to maximize value with two parameterized policies -- a larger expressive base policy trained with a stable imitation learning objective and a light-weight Gaussian edit policy that edits the actions sampled from the base policy toward a higher value distribution. The on-the-fly policy optimizes the actions from the base policy with the learned edit policy and chooses the value maximizing action from the base and edited actions for both sampling and temporal-difference (TD) backup. Our approach yields up to 2-3x improvement in sample efficiency on average over prior methods both in the setting of fine-tuning a pretrained policy given offline data and in leveraging offline data to train online.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 被引用 36 次
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh 等ICLR 2026 · 被引用 13 次
- Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement LearningGuanren Qiao, Ruixiang Ouyang, Sheng Xu, Ruixing Jin 等ICML 2026
- OGPO: Sample Efficient Full-Finetuning of Generative Control PoliciesSarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena 等ICML 2026
它引用的顶会 Paper21
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark 等NeurIPS 2023 · 被引用 296 次
- Efficient Diffusion Policies For Offline Reinforcement LearningBingyi Kang, Xiao Ma, Chao Du, Tianyu Pang 等NeurIPS 2023 · 被引用 195 次
- EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RLSeyed Kamyar Seyed Ghasemipour, Dale Schuurmans, Shixiang Shane GuICML 2021 · 被引用 138 次
相关 Paper
- Flow Matching with Injected Noise for Offline-to-Online Reinforcement LearningYongjae Shin, Jongseong Chae, Jongeui Park, Youngchul SungICLR 2026 · 被引用 1 次
- Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow PoliciesZeyang Li, Sunbochen Tang, Navid AzizanICML 2026 · 被引用 6 次
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Efficient Online Reinforcement Learning for Diffusion PolicyHaitong Ma, Tianyi Chen, Kai Wang, Na Li 等ICML 2025
- Diffusion-based Reinforcement Learning via Q-weighted Variational Policy OptimizationShutong Ding, Ke Hu, Zhenhao Zhang, Kan Ren 等NeurIPS 2024 · 被引用 132 次
