OMPO: A Unified Framework for RL under Policy and Dynamics Shifts
Yu Luo, Tianying Ji, Fuchun Sun, Jianwei Zhang, Huazhe Xu, Xianyuan Zhan
Abstract
Training reinforcement learning policies using environment interaction data collected from varying policies or dynamics presents a fundamental challenge. Existing works often overlook the distribution discrepancies induced by policy or dynamics shifts, or rely on specialized algorithms with task priors, thus often resulting in suboptimal policy performances and high learning variances. In this paper, we identify a unified strategy for online RL policy learning under diverse settings of policy and dynamics shifts: transition occupancy matching. In light of this, we introduce a surrogate policy learning objective by considering the transition occupancy discrepancies and then cast it into a tractable min-max optimization problem through dual reformulation. Our method, dubbed Occupancy-Matching Policy Optimization (OMPO), features a specialized actor-critic structure equipped with a distribution discriminator and a small-size local buffer. We conduct extensive experiments based on the OpenAI Gym, Meta-World, and Panda Robots environments, encompassing policy shifts under stationary and nonstationary dynamics, as well as domain adaption. The results demonstrate that OMPO outperforms the specialized baselines from different categories in all settings. We also find that OMPO exhibits particularly strong performance when combined with domain randomization, highlighting its potential in RL-based robotics applications 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d8e9a22-52be-4452-aa66-72e8a876c285Cited by top-tier papers3
- Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement LearningShangding Gu, Laixi Shi, Muning Wen, Ming Jin et al.ICLR 2025
- Skill Expansion and Composition in Parameter SpaceTenglong Liu, Jianxiong Li, Yinan Zheng, Haoyi Niu et al.ICLR 2025
- A Unified Self-Regulating Training Framework for Federated Deep Reinforcement LearningMeng Xu, Xinhong Chen, Zhongying Chen, Guanyi Zhao et al.AAAI 2026
Builds on17
- Understanding Domain Randomization for Sim-to-real TransferXiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li et al.ICLR 2022 · 164 citations
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee et al.ICML 2020 · 158 citations
- Dropout Q-Functions for Doubly Efficient Reinforcement LearningTakuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi et al.ICLR 2022 · 157 citations
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain ClassifiersBenjamin Eysenbach, Shreyas Chaudhari, Swapnil Asawa, Sergey Levine et al.ICLR 2021 · 120 citations
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li et al.NeurIPS 2022 · 81 citations
Related papers
- SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual DatasetsShenghua Wan, Ziyuan Chen, Le Gan, Shuai Feng et al.ICML 2024 · 1 citation
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu et al.ICML 2024 · 30 citations
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang et al.NeurIPS 2023 · 41 citations
- Cross-Domain Offline Policy Adaptation with Optimal Transport and Dataset ConstraintJiafei Lyu, Mengbei Yan, Zhongjian Qiao, Runze Liu et al.ICLR 2025
- Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement LearningShenzhi Wang, Qisen Yang, Jiawei Gao, Matthieu Gaetan Lin et al.NeurIPS 2023 · 41 citations
