Confounding Robust Deep Reinforcement Learning: A Causal Approach
Mingxuan Li, Junzhe Zhang, Elias Bareinboim
摘要
A key task in Artificial Intelligence is learning effective policies for controlling agents in unknown environments to optimize performance measures. Off-policy learning methods, like Q-learning, allow learners to make optimal decisions based on past experiences. This paper studies off-policy learning from biased data in complex and high-dimensional domains where unobserved confounding cannot be ruled out a priori. Building on the well-celebrated Deep Q-Network (DQN), we propose a novel deep reinforcement learning algorithm robust to confounding biases in observed data. Specifically, our algorithm attempts to find a safe policy for the worst-case environment compatible with the observations. We apply our method to twelve confounded Atari games, and find that it consistently dominates the standard DQN in all games where the observed input to the behavioral and target policies mismatch and unobserved confounders exist.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Causal Flow Q-Learning for Robust Offline Reinforcement LearningMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2026 · 被引用 1 次
- Counterfactual Bootstrap for Robust Meta-Reinforcement LearningAi Bo, Junzhe Zhang, M. Cenk GursoyICML 2026
它引用的顶会 Paper27
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Diffusion for World Modeling: Visual Details Matter in AtariEloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto 等NeurIPS 2024 · 被引用 359 次
相关 Paper
- Off-policy Evaluation for Multiple Actions in the Presence of Unobserved ConfoundersHaolin Wang, Lin Liu, Jiuyong Li, Ziqi Xu 等WWW 2025 · 被引用 1 次
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 5 次
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 被引用 22 次
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 被引用 161 次
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 被引用 10 次
