In-sample Actor Critic for Offline Reinforcement Learning
Hongchang Zhang, Yixiu Mao, Boyuan Wang, Shuncheng He, Yi Xu, Xiangyang Ji
摘要
Offline reinforcement learning suffers from out-of-distribution issue and extrapolation error. Most methods penalize the out-of-distribution state-action pairs or regularize the trained policy towards the behavior policy but cannot guarantee to get rid of extrapolation error. We propose In-sample Actor Critic (IAC) which utilizes sampling-importance resampling to execute in-sample policy evaluation. IAC only uses the target Q-values of the actions in the dataset to evaluate the trained policy, thus avoiding extrapolation error. The proposed method performs unbiased policy evaluation and has a lower variance than importance sampling in many cases. Empirical results show that IAC obtains competitive performance compared to the state-of-the-art methods on Gym-MuJoCo locomotion domains and much more challenging AntMaze domains.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper15
- Supported Value Regularization for Offline Reinforcement LearningYixiu Mao, Hongchang Zhang, Chen Chen, Yi Xu 等NeurIPS 2023 · 被引用 37 次
- Offline Reinforcement Learning with OOD State Correction and OOD Action SuppressionYixiu Mao, Qi Wang, Chen Chen, Yun Qu 等NeurIPS 2024 · 被引用 36 次
- ACT: Empowering Decision Transformer with Dynamic Programming via Advantage ConditioningChenxiao Gao, Chenyang Wu, Mingjun Cao, Rui Kong 等AAAI 2024 · 被引用 31 次
- Doubly Mild Generalization for Offline Reinforcement LearningYixiu Mao, Qi Wang, Yun Qu, Yuhang Jiang 等NeurIPS 2024 · 被引用 30 次
- Constrained Policy Optimization with Explicit Behavior Density For Offline Reinforcement LearningJing Zhang, Chi Zhang, Wenjia Wang, Bingyi JingNeurIPS 2023 · 被引用 19 次
相关 Paper
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang 等AAAI 2025 · 被引用 2 次
- ACTIVE: Offline Reinforcement Learning via Adaptive Imitation and In-sample V-EnsembleTianyuan Chen, Ronglong Cai, Faguo Wu, Xiao ZhangICLR 2025
- Uncertainty Weighted Actor-Critic for Offline Reinforcement LearningYue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind 等ICML 2021 · 被引用 223 次
- The In-Sample Softmax for Offline Reinforcement LearningChenjun Xiao, Han Wang, Yangchen Pan, Adam White 等ICLR 2023 · 被引用 3 次
- Supported Trust Region Optimization for Offline Reinforcement LearningYixiu Mao, Hongchang Zhang, Chen Chen, Yi Xu 等ICML 2023 · 被引用 24 次
