A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware Perspective
Yunpeng Qing, Shunyu Liu, Jingyuan Cong, Kaixuan Chen, Yihe Zhou, Mingli Song
摘要
Offline reinforcement learning endeavors to leverage offline datasets to craft effective agent policy without online interaction, which imposes proper conservative constraints with the support of behavior policies to tackle the out-of-distribution problem. However, existing works often suffer from the constraint conflict issue when offline datasets are collected from multiple behavior policies, i.e., different behavior policies may exhibit inconsistent actions with distinct returns across the state space. To remedy this issue, recent advantage-weighted methods prioritize samples with high advantage values for agent training while inevitably ignoring the diversity of behavior policy. In this paper, we introduce a novel Advantage-Aware Policy Optimization (A2PO) method to explicitly construct advantage-aware policy constraints for offline learning under mixed-quality datasets. Specifically, A2PO employs a conditional variational auto-encoder to disentangle the action distributions of intertwined behavior policies by modeling the advantage values of all training data as conditional variables. Then the agent can follow such disentangled action distribution constraints to optimize the advantage-aware policy towards high advantage values. Extensive experiments conducted on both the single-quality and mixed-quality datasets of the D4RL benchmark demonstrate that A2PO yields results superior to the counterparts. Our code is available at https://github.com/Plankson/A2PO
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Sequential Multi-Agent Dynamic Algorithm ConfigurationChen Lu, Ke Xue, Lei Yuan, Yao Wang 等NeurIPS 2025 · 被引用 8 次
- BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement LearningYunpeng Qing, Yixiao Chi, Shuo Chen, Shunyu Liu 等ICML 2026 · 被引用 4 次
- Cooperative Policy Agreement: Learning Diverse Policy for Offline MARLYihe Zhou, Yuxuan Zheng, Yue Hu, Kaixuan Chen 等AAAI 2025 · 被引用 2 次
- Tackling Continual Offline RL through Selective Weights Activation on Aligned SpacesJifeng Hu, Sili Huang, Li Shen, Zhejian Yang 等NeurIPS 2025 · 被引用 2 次
- Direct Flow Q-LearningShicheng Cao, Jingrui Jia, Wenyu Li, Feng Duan 等ICML 2026
它引用的顶会 Paper30
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
相关 Paper
- Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningTenglong Liu, Yang Li, Yixing Lan, Hao Gao 等ICML 2024 · 被引用 15 次
- Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement LearningPrajwal Koirala, Zhanhong Jiang, Soumik Sarkar, Cody H. FlemingICLR 2025
- LAPO: Latent-Variable Advantage-Weighted Policy Optimization for Offline Reinforcement LearningXi Chen, Ali Ghadirzadeh, Tianhe Yu, Jianhao Wang 等NeurIPS 2022 · 被引用 52 次
- Weighted Policy Constraints for Offline Reinforcement LearningZhiyong Peng, Changlin Han, Yadong Liu, Zongtan ZhouAAAI 2023 · 被引用 18 次
- SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual DatasetsShenghua Wan, Ziyuan Chen, Le Gan, Shuai Feng 等ICML 2024 · 被引用 1 次
