Provably Efficient Causal Reinforcement Learning with Confounded Observational Data
Lingxiao Wang, Zhuoran Yang, Zhaoran Wang
Abstract
Empowered by expressive function approximators such as neural networks, deep reinforcement learning (DRL) achieves tremendous empirical successes. However, learning expressive function approximators requires collecting a large dataset (interventional data) by interacting with the environment. Such a lack of sample efficiency prohibits the application of DRL to critical scenarios, e.g., autonomous driving and personalized medicine, since trial and error in the online setting is often unsafe and even unethical. In this paper, we study how to incorporate the dataset (observational data) collected offline, which is often abundantly available in practice, to improve the sample efficiency in the online setting. To incorporate the observational data, we face two challenges. (a) The behavior policy that generates the observational data may depend on possibly unobserved random variables (confounders), which at the same time, affect the received rewards and transition dynamics. Such a confounding issue makes the observational data uninformative and even misleading for decision making in the online setting. (b) Exploration in the online setting requires quantifying the uncertainty that remains given both the observational and interventional data. In particular, it remains unclear how to quantify the amount of information carried over by the confounded observational data, which plays a key role in constructing the bonus and characterizing the regret. To address the two challenges, we propose the deconfounded optimistic value iteration (DOVI) algorithm, which incorporates the confounded observational data in a provably efficient manner. More specifically, DOVI explicitly adjusts for the confounding bias in the observational data, where the confounders are partially observed or unobserved. In both cases, such adjustments allow us to construct the bonus based on a notion of information gain, which takes into account the amount of information acquired from the offline setting. In particular, we prove that the regret of DOVI is smaller than the optimal regret achievable in the pure online setting by a multiplicative factor, which decreases towards zero when the confounded observational data are more informative upon the adjustments. Our algorithm and analysis serve as a step towards causal reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b1d80c5-128a-41dc-aa06-39c142d85a19Cited by top-tier papers17
- Seeing is not Believing: Robust Reinforcement Learning against Spurious CorrelationWenhao Ding, Laixi Shi, Yuejie Chi, Ding ZhaoNeurIPS 2023 · 39 citations
- Dynamic Bottleneck for Robust Self-Supervised ExplorationChenjia Bai, Lingxiao Wang, Lei Han, Animesh Garg et al.NeurIPS 2021 · 36 citations
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov et al.NeurIPS 2023 · 31 citations
- A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision ProcessesChengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan JiangICML 2022 · 31 citations
- A Relational Intervention Approach for Unsupervised Dynamics Generalization in Model-Based Reinforcement LearningJiaxian Guo, Mingming Gong, Dacheng TaoICLR 2022 · 21 citations
Builds on2
Related papers
- Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision ProcessesMiao Lu, Yifei Min, Zhaoran Wang, Zhuoran YangICLR 2023
- Provably Efficient Offline Reinforcement Learning for Partially Observable Markov Decision ProcessesHongyi Guo, Qi Cai, Yufeng Zhang, Zhuoran Yang et al.ICML 2022 · 17 citations
- Delphic Offline Reinforcement Learning under Nonidentifiable Hidden ConfoundingAlizée Pace, Hugo Yèche, Bernhard Schölkopf, Gunnar Rätsch et al.ICLR 2024 · 9 citations
- On Covariate Shift of Latent Confounders in Imitation and Reinforcement LearningGuy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor et al.ICLR 2022 · 16 citations
- Designing Optimal Dynamic Treatment Regimes: A Causal Reinforcement Learning ApproachJunzhe ZhangICML 2020 · 78 citations
