Recovering from Out-of-sample States via Inverse Dynamics in Offline Reinforcement Learning
Ke Jiang, Jia-Yu Yao, Xiaoyang Tan
Abstract
We deal with the state distributional shift problem commonly encountered in offline reinforcement learning during test, where the agent tends to take unreliable actions at out-of-sample (unseen) states. Our idea is to encourage the agent to follow the so called state recovery principle when taking actions, i.e., besides long-term return, the immediate consequences of the current action should also be taken into account and those capable of recovering the state distribution of the behavior policy are preferred. For this purpose, an inverse dynamics model is learned and employed to guide the state recovery behavior of the new policy. Theoretically, we show that the proposed method helps aligning the transited state distribution of the new policy with the offline dataset at out-of-sample states, without the need of explicitly predicting the transited state distribution, which is usually difficult in high-dimensional and complicated environments. The effectiveness and feasibility of the proposed method is demonstrated with the state-of-the-art performance on the general offline RL benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c69c93a2-52bc-48fd-9a75-4d9de2b173f8Cited by top-tier papers2
- Offline Reinforcement Learning with OOD State Correction and OOD Action SuppressionYixiu Mao, Qi Wang, Chen Chen, Yun Qu et al.NeurIPS 2024 · 36 citations
- Variational OOD State Correction for Offline Reinforcement LearningKe Jiang, Wen Jiang, Xiaoyang TanAAAI 2026
Builds on10
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 430 citations
Related papers
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang et al.AAAI 2025 · 2 citations
- Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement LearningShentao Yang, Yihao Feng, Shujian Zhang, Mingyuan ZhouICML 2022 · 14 citations
- State Deviation Correction for Offline Reinforcement LearningHongchang Zhang, Jianzhun Shao, Yuhang Jiang, Shuncheng He et al.AAAI 2022 · 18 citations
- CLARE: Conservative Model-Based Reward Learning for Offline Inverse Reinforcement LearningSheng Yue, Guanbo Wang, Wei Shao, Zhaofeng Zhang et al.ICLR 2023 · 6 citations
- S2P: State-conditioned Image Synthesis for Data Augmentation in Offline Reinforcement LearningDaesol Cho, Dongseok Shim, H. Jin KimNeurIPS 2022 · 14 citations
