Doubly-Robust Estimation of Counterfactual Policy Mean Embeddings
Houssam Zenati, Bariscan Bozkurt, Arthur Gretton
摘要
Estimating the distribution of outcomes under counterfactual policies is critical for decision-making in domains such as recommendation, advertising, and healthcare. We propose and analyze a novel framework-Counterfactual Policy Mean Embedding (CPME)-that represents the entire counterfactual outcome distribution in a reproducing kernel Hilbert space (RKHS), enabling flexible and nonparametric distributional off-policy evaluation. We introduce both a plug-in estimator and a doubly robust estimator; the latter enjoys improved convergence rates by correcting for bias in both the outcome embedding and propensity models. Building on this, we develop a doubly robust kernel test statistic for hypothesis testing, which achieves asymptotic normality and thus enables computationally efficient testing and straightforward construction of confidence intervals. Our framework also supports sampling from the counterfactual distribution. Numerical simulations illustrate the practical benefits of CPME over existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- A Measure-Theoretic Approach to Kernel Conditional Mean EmbeddingsJunhyung Park, Krikamol MuandetNeurIPS 2020 · 被引用 123 次
- Optimal Rates for Regularized Conditional Mean Embedding LearningZhu Li, Dimitri Meunier, Mattes Mollenhauer, Arthur GrettonNeurIPS 2022 · 被引用 69 次
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller 等NeurIPS 2021 · 被引用 64 次
- Conditional Distributional Treatment Effect with Kernel Conditional Mean Embeddings and U-Statistic RegressionJunhyung Park, Uri Shalit, Bernhard Schölkopf, Krikamol MuandetICML 2021 · 被引用 46 次
- Off-Policy Risk Assessment in Contextual BanditsAudrey Huang, Liu Leqi, Zachary C. Lipton, Kamyar AzizzadenesheliNeurIPS 2021 · 被引用 44 次
相关 Paper
- Counterfactual Density Estimation using Kernel Stein DiscrepanciesDiego Martinez-Taboada, Edward KennedyICLR 2024 · 被引用 8 次
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 被引用 60 次
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 被引用 1 次
- Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningNathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan ZhouICML 2022 · 被引用 39 次
- Accountable Off-Policy Evaluation With Kernel Bellman StatisticsYihao Feng, Tongzheng Ren, Ziyang Tang, Qiang LiuICML 2020 · 被引用 45 次
