Off-policy estimation with adaptively collected data: the power of online learning
Jeonghwan Lee, Cong Ma
摘要
We consider estimation of a linear functional of the treatment effect using adaptively collected data. This task finds a variety of applications including the off-policy evaluation (OPE) in contextual bandits, and estimation of the average treatment effect (ATE) in causal inference. While a certain class of augmented inverse propensity weighting (AIPW) estimators enjoys desirable asymptotic properties including the semi-parametric efficiency, much less is known about their non-asymptotic theory with adaptively collected data. To fill in the gap, we first establish generic upper bounds on the mean-squared error of the class of AIPW estimators that crucially depends on a sequentially weighted error between the treatment effect and its estimates. Motivated by this, we also propose a general reduction scheme that allows one to produce a sequence of estimates for the treatment effect via online learning to minimize the sequentially weighted estimation error. To illustrate this, we provide three concrete instantiations in (1) the tabular case; (2) the case of linear function approximation; and (3) the case of general function approximation for the outcome model. We then provide a local minimax lower bound to show the instance-dependent optimality of the AIPW estimator using no-regret online learning algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Pessimistic Data Integration for Policy EvaluationXiangkun Wu, Ting Li, Gholamali Aminian, Armin Behnamnia 等NeurIPS 2025 · 被引用 2 次
- Efficient Adaptive Experimentation with NoncomplianceMiruna Oprescu, Brian Cho, Nathan KallusNeurIPS 2025
它引用的顶会 Paper11
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 被引用 115 次
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
- Optimal Off-Policy Evaluation from Multiple Logging PoliciesNathan Kallus, Yuta Saito, Masatoshi UeharaICML 2021 · 被引用 44 次
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 被引用 34 次
相关 Paper
- Optimistic Algorithms for Adaptive Estimation of the Average Treatment EffectOjash Neopane, Aaditya Ramdas, Aarti SinghICML 2025
- Marginal Density Ratio for Off-Policy Evaluation in Contextual BanditsMuhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish, Jean-Francois TonNeurIPS 2023 · 被引用 14 次
- PUATE: Efficient ATE Estimation from Treated (Positive) and Unlabeled UnitsMasahiro Kato, Fumiaki Kozai, Ryo InokuchiNeurIPS 2025
- Online Multi-Armed Bandits with Adaptive InferenceMaria Dimakopoulou, Zhimei Ren, Zhengyuan ZhouNeurIPS 2021 · 被引用 47 次
- Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate ChoiceMasahiro Kato, Akihiro Oga, Wataru Komatsubara, Ryo InokuchiICML 2024 · 被引用 12 次
