Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning
Shuguang Yu, Shuxing Fang, Ruixin Peng, Zhengling Qi, Fan Zhou, Chengchun Shi
Abstract
This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured confounding assumption to model the system dynamics in causal reinforcement learning and develop a two-way deconfounder algorithm that devises a neural tensor network to simultaneously learn both the unmeasured confounders and the system dynamics, based on which a model-based estimator can be constructed for consistent policy value estimation. We illustrate the effectiveness of the proposed estimator through theoretical results and numerical experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d91db35-aaec-43e2-bec1-eb81159b9de4Cited by top-tier papers3
- Pessimistic Data Integration for Policy EvaluationXiangkun Wu, Ting Li, Gholamali Aminian, Armin Behnamnia et al.NeurIPS 2025 · 2 citations
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Unraveling the Interplay between Carryover Effects and Reward Autocorrelations in Switchback ExperimentsQianglin Wen, Chengchun Shi, Ying Yang, Niansheng Tang et al.ICML 2025
Builds on24
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- Time Series Deconfounder: Estimating Treatment Effects over Time in the Presence of Hidden ConfoundersIoana Bica, Ahmed M. Alaa, Mihaela van der SchaarICML 2020 · 133 citations
- Provably Good Batch Off-Policy Reinforcement Learning Without Great ExplorationYao Liu, Adith Swaminathan, Alekh Agarwal, Emma BrunskillNeurIPS 2020 · 97 citations
Related papers
- Off-policy Evaluation for Multiple Actions in the Presence of Unobserved ConfoundersHaolin Wang, Lin Liu, Jiuyong Li, Ziqi Xu et al.WWW 2025 · 1 citation
- Model-Free and Model-Based Policy Evaluation when Causality is UncertainDavid Bruns-SmithICML 2021 · 14 citations
- Model-based Reinforcement Learning for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICML 2024 · 5 citations
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo et al.ICML 2023 · 24 citations
- Deep Proxy Causal Learning and its Application to Confounded Bandit Policy EvaluationLiyuan Xu, Heishiro Kanagawa, Arthur GrettonNeurIPS 2021 · 52 citations
