Delphic Offline Reinforcement Learning under Nonidentifiable Hidden Confounding
Alizée Pace, Hugo Yèche, Bernhard Schölkopf, Gunnar Rätsch, Guy Tennenholtz
Abstract
A prominent challenge of offline reinforcement learning (RL) is the issue of hidden confounding: unobserved variables may influence both the actions taken by the agent and the observed outcomes. Hidden confounding can compromise the validity of any causal conclusion drawn from data and presents a major obstacle to effective offline RL. In the present paper, we tackle the problem of hidden confounding in the nonidentifiable setting. We propose a definition of uncertainty due to hidden confounding bias, termed delphic uncertainty, which uses variation over world models compatible with the observations, and differentiate it from the well-known epistemic and aleatoric uncertainties. We derive a practical method for estimating the three types of uncertainties, and construct a pessimistic offline RL algorithm to account for them. Our method does not assume identifiability of the unobserved confounders, and attempts to reduce the amount of confounding bias. We demonstrate through extensive experiments and ablations the efficacy of our approach on a sepsis management benchmark, as well as on electronic health records. Our results suggest that nonidentifiable hidden confounding bias can be mitigated to improve offline RL solutions in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3e207d8-c821-4831-a11e-0daf1a14d404Cited by top-tier papers5
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 5 citations
- BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement LearningHaohong Lin, Wenhao Ding, Jian Chen, Laixi Shi et al.NeurIPS 2024 · 5 citations
- Learning Decision Policies with Instrumental Variables through Double Machine LearningDaqian Shao, Ashkan Soleymani, Francesco Quinzan, Marta KwiatkowskaICML 2024 · 4 citations
- Identifying Latent State-Transition Processes for Individualized Reinforcement LearningYuewen Sun, Biwei Huang, Yu Yao, Donghuo Zeng et al.NeurIPS 2024 · 2 citations
- Turning Sand to Gold: Recycling Data to Bridge On-Policy and Off-Policy Learning via Causal BoundTal Fiskus, Uri ShahamNeurIPS 2025
Builds on24
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
Related papers
- Provably Efficient Causal Reinforcement Learning with Confounded Observational DataLingxiao Wang, Zhuoran Yang, Zhaoran WangNeurIPS 2021 · 61 citations
- Provably Efficient Offline Reinforcement Learning for Partially Observable Markov Decision ProcessesHongyi Guo, Qi Cai, Yufeng Zhang, Zhuoran Yang et al.ICML 2022 · 17 citations
- Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationMuchan Tao, Haonan Qin, Yuqi Fang, Caifeng Shan et al.AAAI 2026
- Confounding-Robust Policy Evaluation in Infinite-Horizon Reinforcement LearningNathan Kallus, Angela ZhouNeurIPS 2020 · 78 citations
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 81 citations
