Off-policy estimation with adaptively collected data: the power of online learning
Jeonghwan Lee, Cong Ma
Abstract
We consider estimation of a linear functional of the treatment effect using adaptively collected data. This task finds a variety of applications including the off-policy evaluation (OPE) in contextual bandits, and estimation of the average treatment effect (ATE) in causal inference. While a certain class of augmented inverse propensity weighting (AIPW) estimators enjoys desirable asymptotic properties including the semi-parametric efficiency, much less is known about their non-asymptotic theory with adaptively collected data. To fill in the gap, we first establish generic upper bounds on the mean-squared error of the class of AIPW estimators that crucially depends on a sequentially weighted error between the treatment effect and its estimates. Motivated by this, we also propose a general reduction scheme that allows one to produce a sequence of estimates for the treatment effect via online learning to minimize the sequentially weighted estimation error. To illustrate this, we provide three concrete instantiations in (1) the tabular case; (2) the case of linear function approximation; and (3) the case of general function approximation for the outcome model. We then provide a local minimax lower bound to show the instance-dependent optimality of the AIPW estimator using no-regret online learning algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bf52d4f-c29f-4923-9028-cce990f8e2aaCited by top-tier papers2
- Pessimistic Data Integration for Policy EvaluationXiangkun Wu, Ting Li, Gholamali Aminian, Armin Behnamnia et al.NeurIPS 2025 · 2 citations
- Efficient Adaptive Experimentation with NoncomplianceMiruna Oprescu, Brian Cho, Nathan KallusNeurIPS 2025
Builds on11
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 115 citations
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
- Optimal Off-Policy Evaluation from Multiple Logging PoliciesNathan Kallus, Yuta Saito, Masatoshi UeharaICML 2021 · 44 citations
- Off-Policy Evaluation for Large Action Spaces via Conjunct Effect ModelingYuta Saito, Qingyang Ren, Thorsten JoachimsICML 2023 · 34 citations
Related papers
- Optimistic Algorithms for Adaptive Estimation of the Average Treatment EffectOjash Neopane, Aaditya Ramdas, Aarti SinghICML 2025
- Marginal Density Ratio for Off-Policy Evaluation in Contextual BanditsMuhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish, Jean-Francois TonNeurIPS 2023 · 14 citations
- PUATE: Efficient ATE Estimation from Treated (Positive) and Unlabeled UnitsMasahiro Kato, Fumiaki Kozai, Ryo InokuchiNeurIPS 2025
- Online Multi-Armed Bandits with Adaptive InferenceMaria Dimakopoulou, Zhimei Ren, Zhengyuan ZhouNeurIPS 2021 · 47 citations
- Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate ChoiceMasahiro Kato, Akihiro Oga, Wataru Komatsubara, Ryo InokuchiICML 2024 · 12 citations
