EXPO: Stable Reinforcement Learning with Expressive Policies
Perry Dong, Qiyang Li, Dorsa Sadigh, Chelsea Finn
Abstract
We study the problem of training and fine-tuning expressive policies with online reinforcement learning (RL) given an offline dataset. Training expressive policy classes with online RL present a unique challenge of stable value maximization. Unlike simpler Gaussian policies commonly used in online RL, expressive policies like diffusion and flow-matching policies are parameterized by a long denoising chain, which hinders stable gradient propagation from actions to policy parameters when optimizing against some value function. Our key insight is that we can address stable value maximization by avoiding direct optimization over value with the expressive policy and instead construct an on-the-fly RL policy to maximize Q-value. We propose Expressive Policy Optimization (EXPO), a sample-efficient online RL algorithm that utilizes an on-the-fly policy to maximize value with two parameterized policies -- a larger expressive base policy trained with a stable imitation learning objective and a light-weight Gaussian edit policy that edits the actions sampled from the base policy toward a higher value distribution. The on-the-fly policy optimizes the actions from the base policy with the learned edit policy and chooses the value maximizing action from the base and edited actions for both sampling and temporal-difference (TD) backup. Our approach yields up to 2-3x improvement in sample efficiency on average over prior methods both in the setting of fine-tuning a pretrained policy given offline data and in leveraging offline data to train online.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f887647-2977-41f4-936f-47746a09fbb3Cited by top-tier papers4
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 36 citations
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh et al.ICLR 2026 · 13 citations
- Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement LearningGuanren Qiao, Ruixiang Ouyang, Sheng Xu, Ruixing Jin et al.ICML 2026
- OGPO: Sample Efficient Full-Finetuning of Generative Control PoliciesSarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena et al.ICML 2026
Builds on21
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
- Efficient Diffusion Policies For Offline Reinforcement LearningBingyi Kang, Xiao Ma, Chao Du, Tianyu Pang et al.NeurIPS 2023 · 195 citations
- EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RLSeyed Kamyar Seyed Ghasemipour, Dale Schuurmans, Shixiang Shane GuICML 2021 · 138 citations
Related papers
- Flow Matching with Injected Noise for Offline-to-Online Reinforcement LearningYongjae Shin, Jongseong Chae, Jongeui Park, Youngchul SungICLR 2026 · 1 citation
- Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow PoliciesZeyang Li, Sunbochen Tang, Navid AzizanICML 2026 · 6 citations
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Efficient Online Reinforcement Learning for Diffusion PolicyHaitong Ma, Tianyi Chen, Kai Wang, Na Li et al.ICML 2025
- Diffusion-based Reinforcement Learning via Q-weighted Variational Policy OptimizationShutong Ding, Ke Hu, Zhenhao Zhang, Kan Ren et al.NeurIPS 2024 · 132 citations
