Variational OOD State Correction for Offline Reinforcement Learning
Ke Jiang, Wen Jiang, Xiaoyang Tan
摘要
The performance of Offline reinforcement learning is significantly impacted by the issue of state distributional shift, and out-of-distribution (OOD) state correction is a popular approach to address this problem. In this paper, we propose a novel method named Density-Aware Safety Perception (DASP) for OOD state correction. Specifically, our method encourages the agent to prioritize actions that lead to outcomes with higher data density, thereby promoting its operation within or the return to in-distribution (safe) regions. To achieve this, we optimize the objective within a variational framework that concurrently considers both the potential outcomes of decision-making and their density, thus providing crucial contextual information for safe decision-making. Finally, we validate the effectiveness and feasibility of our proposed method through extensive experimental evaluations on the offline MuJoCo and AntMaze suites. ...... OOD state State deviation DASP-based OOD state correction High-density regions (dataset) Low density (OOD) High density (In-distribution) Unrecoverable OOD state Agent High-density state
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng 等ICLR 2022 · 被引用 173 次
- Lyapunov Density Models: Constraining Distribution Shift in Learning-Based ControlKatie Kang, Paula Gradu, Jason J. Choi, Michael Janner 等ICML 2022 · 被引用 39 次
- Constrained Policy Optimization with Explicit Behavior Density For Offline Reinforcement LearningJing Zhang, Chi Zhang, Wenjia Wang, Bingyi JingNeurIPS 2023 · 被引用 19 次
- State Deviation Correction for Offline Reinforcement LearningHongchang Zhang, Jianzhun Shao, Yuhang Jiang, Shuncheng He 等AAAI 2022 · 被引用 18 次
相关 Paper
- Offline Reinforcement Learning with OOD State Correction and OOD Action SuppressionYixiu Mao, Qi Wang, Chen Chen, Yun Qu 等NeurIPS 2024 · 被引用 36 次
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang 等AAAI 2025 · 被引用 2 次
- Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency modelJing Zhang, Linjiajie Fang, Kexin Shi, Wenjia Wang 等NeurIPS 2024 · 被引用 14 次
- Learning from Sparse Offline Datasets via Conservative Density EstimationZhepeng Cen, Zuxin Liu, Zitong Wang, Yihang Yao 等ICLR 2024 · 被引用 12 次
- VOCE: Variational Optimization with Conservative Estimation for Offline Safe Reinforcement LearningJiayi Guan, Guang Chen, Jiaming Ji, Long Yang 等NeurIPS 2023 · 被引用 19 次
