Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery
Yiqin Yang, Hao Hu, Wenzhe Li, Siyuan Li, Jun Yang, Qianchuan Zhao, Chongjie Zhang
Abstract
Offline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe. Previous works have shown that extracting primitive skills from the recurring and temporally extended structures in the logged data yields better learning. However, these methods suffer greatly when the primitives have limited representation ability to recover the original policy space, especially in offline settings. In this paper, we give a quantitative characterization of the performance of offline hierarchical learning and highlight the importance of learning lossless primitives. To this end, we propose to use a flow-based structure as the representation for low-level policies. This allows us to represent the behaviors in the dataset faithfully while keeping the expression ability to recover the whole policy space. We show that such lossless primitives can drastically improve the performance of hierarchical policies. The experimental results and extensive ablation studies on the standard D4RL benchmark show that our method has a good representation ability for policies and achieves superior performance in most tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c99380e-c567-4fd9-9fb9-e32cf783b5e1Cited by top-tier papers8
- Hierarchical Diffusion for Offline Decision MakingWenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan ZhaICML 2023 · 80 citations
- Unsupervised Behavior Extraction via Random Intent PriorsHao Hu, Yiqin Yang, Jianing Ye, Ziqing Mai et al.NeurIPS 2023 · 15 citations
- Efficient Planning with Latent DiffusionWenhao LiICLR 2024 · 15 citations
- Bayesian Design Principles for Offline-to-Online Reinforcement LearningHao Hu, Yiqin Yang, Jianing Ye, Chengjie Wu et al.ICML 2024 · 10 citations
- Planning, Fast and Slow: Online Reinforcement Learning with Action-Free Offline Data via Multiscale PlannersChengjie Wu, Hao Hu, Yiqin Yang, Ning Zhang et al.ICML 2024 · 3 citations
Builds on19
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
Related papers
- Offline Hierarchical Reinforcement Learning via Inverse OptimizationCarolin Schmidt, Daniele Gammelli, James Harrison, Marco Pavone et al.ICLR 2025
- OPAL: Offline Primitive Discovery for Accelerating Offline Reinforcement LearningAnurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine et al.ICLR 2021 · 15 citations
- Action-Sufficient Goal RepresentationsJinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup MoonICML 2026 · 1 citation
- Reasoning with Latent Diffusion in Offline Reinforcement LearningSiddarth Venkatraman, Shivesh Khaitan, Ravi Tej Akella, John M. Dolan et al.ICLR 2024 · 39 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
