Adversarial Counterfactual Environment Model Learning
Xiong-Hui Chen, Yang Yu, Zhengmao Zhu, Zhihua Yu, Zhenjun Chen, Chenghe Wang, Yinan Wu, Rong-Jun Qin, Hongqiu Wu, Ruijin Ding, Fangsheng Huang
Abstract
A good model for action-effect prediction, named environment model, is important to achieve sample-efficient decision-making policy learning in many domains like robot control, recommender systems, and patients’ treatment selection. We can take unlimited trials with such a model to identify the appropriate actions so that the costs of queries in the real world can be saved. It requires the model to correctly handle unseen data, also called counterfactual data. However, standard data fitting techniques do not automatically achieve such generalization ability and commonly result in unreliable models. In this work, we introduce counterfactual-query risk minimization (CQRM) in model learning for generalizing to a counterfactual dataset queried by a specific target policy. Since the target policies can be various and unknown in policy learning, we propose an adversarial CQRM objective in which the model learns on counterfactual data queried by adversarial policies, and finally derive a tractable solution GALILEO. We also discover that adversarial CQRM is closely related to the adversarial model learning, explaining the effectiveness of the latter. We apply GALILEO in synthetic tasks and a real-world application. The results show that GALILEO makes accurate predictions on counterfactual data and thus significantly improves policies in real-world testing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Adversarial Model for Offline Reinforcement LearningMohak Bhardwaj, Tengyang Xie, Byron Boots, Nan Jiang et al.NeurIPS 2023 · 44 citations
- Plan To Predict: Learning an Uncertainty-Foreseeing Model For Model-Based Reinforcement LearningZifan Wu, Chao Yu, Chen Chen, Jianye Hao et al.NeurIPS 2022 · 28 citations
- Models as Agents: Optimizing Multi-Step Predictions of Interactive Local Models in Model-Based Multi-Agent Reinforcement LearningZifan Wu, Chao Yu, Chen Chen, Jianye Hao et al.AAAI 2023 · 16 citations
- Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement LearningXu-Hui Liu, Tian-Shuo Liu, Shengyi Jiang, Ruifeng Chen et al.ICML 2024 · 10 citations
- Policy-conditioned Environment Models are More GeneralizableRuifeng Chen, Xiong-Hui Chen, Yihao Sun, Siyuan Xiao et al.ICML 2024 · 1 citation
Builds on14
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- Estimating the Effects of Continuous-valued Interventions using Generative Adversarial NetworksIoana Bica, James Jordon, Mihaela van der SchaarNeurIPS 2020 · 137 citations
- Counterfactual Data Augmentation using Locally Factored DynamicsSilviu Pitis, Elliot Creager, Animesh GargNeurIPS 2020 · 126 citations
Related papers
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 24 citations
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 130 citations
- ACAMDA: Improving Data Efficiency in Reinforcement Learning through Guided Counterfactual Data AugmentationYuewen Sun, Erli Wang, Biwei Huang, Chaochao Lu et al.AAAI 2024
- Distributionally Robust Counterfactual Risk MinimizationLouis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova et al.AAAI 2020 · 48 citations
- Continuous-Time Counterfactual Quantile Learning for Risk-Sensitive Policy OptimizationYi He, Anpeng Wu, Ruoxuan Xiong, Yingrong Wang et al.KDD 2026
