The Adaptive Doubly Robust Estimator and a Paradox Concerning Logging Policy
Masahiro Kato, Kenichiro McAlinn, Shota Yasui
Abstract
The doubly robust (DR) estimator, which consists of two nuisance parameters, the conditional mean outcome and the logging policy (the probability of choosing an action), is crucial in causal inference. This paper proposes a DR estimator for dependent samples obtained from adaptive experiments. To obtain an asymptotically normal semiparametric estimator from dependent samples with non-Donsker nuisance estimators, we propose adaptive-fitting as a variant of sample-splitting. We also report an empirical paradox that our proposed DR estimator tends to show better performances compared to other estimators utilizing the true logging policy. While a similar phenomenon is known for estimators with i.i.d. samples, traditional explanations based on asymptotic efficiency cannot elucidate our case with dependent samples. We confirm this hypothesis through simulation studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate ChoiceMasahiro Kato, Akihiro Oga, Wataru Komatsubara, Ryo InokuchiICML 2024 · 12 citations
- Causal-EPIG: Causally Aligned Active CATE EstimationErdun Gao, Jake Fawkes, Dino SejdinovicICML 2026 · 3 citations
- Observationally Informed Adaptive Causal Experimental DesignErdun Gao, Liang Zhang, Jake Fawkes, Aoqi Zuo et al.KDD 2026 · 1 citation
- Efficient Adaptive Experimentation with NoncomplianceMiruna Oprescu, Brian Cho, Nathan KallusNeurIPS 2025
- PUATE: Efficient ATE Estimation from Treated (Positive) and Unlabeled UnitsMasahiro Kato, Fumiaki Kozai, Ryo InokuchiNeurIPS 2025
Builds on6
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 329 citations
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 115 citations
- Estimating Identifiable Causal Effects through Double Machine LearningYonghan Jung, Jin Tian, Elias BareinboimAAAI 2021 · 70 citations
- Statistical Inference with M-Estimators on Adaptively Collected DataKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2021 · 66 citations
Related papers
- Off-Policy Evaluation via Adaptive Weighting with Data from Contextual BanditsRuohan Zhan, Vitor Hadad, David A. Hirshberg, Susan AtheyKDD 2021 · 22 citations
- Continuous Treatment Effects with Surrogate OutcomesZhenghao Zeng, David Arbour, Avi Feller, Raghavendra Addanki et al.ICML 2024 · 4 citations
- Online Multi-Armed Bandits with Adaptive InferenceMaria Dimakopoulou, Zhimei Ren, Zhengyuan ZhouNeurIPS 2021 · 47 citations
- Doubly Robust Causal Effect Estimation under Networked Interference via Targeted LearningWeilin Chen, Ruichu Cai, Zeqin Yang, Jie Qiao et al.ICML 2024 · 17 citations
- Distributionally Robust Policy Learning under Concept DriftsJingyuan Wang, Zhimei Ren, Ruohan Zhan, Zhengyuan ZhouICML 2025
