Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning
Shuguang Yu, Shuxing Fang, Ruixin Peng, Zhengling Qi, Fan Zhou, Chengchun Shi
摘要
This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured confounding assumption to model the system dynamics in causal reinforcement learning and develop a two-way deconfounder algorithm that devises a neural tensor network to simultaneously learn both the unmeasured confounders and the system dynamics, based on which a model-based estimator can be constructed for consistent policy value estimation. We illustrate the effectiveness of the proposed estimator through theoretical results and numerical experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pessimistic Data Integration for Policy EvaluationXiangkun Wu, Ting Li, Gholamali Aminian, Armin Behnamnia 等NeurIPS 2025 · 被引用 2 次
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Unraveling the Interplay between Carryover Effects and Reward Autocorrelations in Switchback ExperimentsQianglin Wen, Chengchun Shi, Ying Yang, Niansheng Tang 等ICML 2025
它引用的顶会 Paper24
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- Time Series Deconfounder: Estimating Treatment Effects over Time in the Presence of Hidden ConfoundersIoana Bica, Ahmed M. Alaa, Mihaela van der SchaarICML 2020 · 被引用 133 次
- Provably Good Batch Off-Policy Reinforcement Learning Without Great ExplorationYao Liu, Adith Swaminathan, Alekh Agarwal, Emma BrunskillNeurIPS 2020 · 被引用 97 次
相关 Paper
- Off-policy Evaluation for Multiple Actions in the Presence of Unobserved ConfoundersHaolin Wang, Lin Liu, Jiuyong Li, Ziqi Xu 等WWW 2025 · 被引用 1 次
- Model-Free and Model-Based Policy Evaluation when Causality is UncertainDavid Bruns-SmithICML 2021 · 被引用 14 次
- Model-based Reinforcement Learning for Confounded POMDPsMao Hong, Zhengling Qi, Yanxun XuICML 2024 · 被引用 5 次
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo 等ICML 2023 · 被引用 24 次
- Deep Proxy Causal Learning and its Application to Confounded Bandit Policy EvaluationLiyuan Xu, Heishiro Kanagawa, Arthur GrettonNeurIPS 2021 · 被引用 52 次
