Efficient Adaptive Experimentation with Noncompliance
Miruna Oprescu, Brian Cho, Nathan Kallus
Abstract
We study the problem of estimating the average treatment effect (ATE) in adaptive experiments where treatment can only be encouraged-rather than directly assigned-via a binary instrumental variable. Building on semiparametric efficiency theory, we derive the efficiency bound for ATE estimation under arbitrary, history-dependent instrument-assignment policies, and show it is minimized by a variance-aware allocation rule that balances outcome noise and compliance variability. Leveraging this insight, we introduce AMRIV-an Adaptive, Multiply-Robust estimator for Instrumental-Variable settings with variance-optimal assignment. AMRIV pairs (i) an online policy that adaptively approximates the optimal allocation with (ii) a sequential, influence-function-based estimator that attains the semiparametric efficiency bound while retaining multiply-robust consistency. We establish asymptotic normality, explicit convergence rates, and anytime-valid asymptotic confidence sequences that enable sequential inference. Finally, we demonstrate the practical effectiveness of our approach through empirical studies, showing that adaptive instrument assignment, when combined with the AMRIV estimator, yields improved efficiency and robustness compared to existing baselines.
In such applications, reducing estimator variance is not merely a statistical preference; it determines how quickly and confidently a study can reach conclusions. Because high variance delays both 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
the detection of harmful effects and the confirmation of beneficial ones, adaptive designs that learn to allocate instruments in variance-minimizing ways can mitigate these risks, yielding tighter confidence sequences and enabling earlier, statistically valid stopping decisions. Despite extensive work on adaptive designs for settings where treatment can be directly enforced [13,23,27], adaptive experimentation under noncompliance-where treatment is voluntary but encouragement can be adaptively controlled, remains largely unexplored, even though it describes many real-world scenarios. This paper addresses this gap. We study the problem of estimating the average treatment effect (ATE) in a sequential experiment where the experimenter can assign only a binary instrument, while the treatment itself is determined endogenously. Building on the semiparametric framework of Wang and Tchetgen Tchetgen [54], which identifies the ATE under an unconfounded compliance assumption and provides a multiply robust, efficient influence-function-based estimator, we extend this framework to the adaptive setting. Specifically:
• We derive the semiparametric efficiency bound and characterize the variance-optimal adaptive policy that minimizes this bound through a covariate-dependent instrument assignment.
• We introduce AMRIV, an Adaptive, Multiply Robust estimator for IV settings, which applies a sequential, plug-in version of the efficient influence function evaluated under the adaptive policy.
• We establish strong theoretical guarantees, including asymptotic normality, explicit convergence rates, multiply robust consistency, and time-uniform asymptotic confidence sequences for valid inference at arbitrary stopping times.
• We demonstrate practical effectiveness in both synthetic and semi-synthetic studies, showing improved efficiency and robustness over non-adaptive baselines and alternatives. In contrast to prior work on adaptive design with instruments [3,9,22,61], our method focuses on point estimation of the ATE, achieves semiparametric efficiency, and supports multiply robust inference under adaptive assignment. To our knowledge, this is the first method to bring the full suite of modern semiparametric tools-efficient influence functions, adaptive policy learning, robust plug-in estimation, and anytime-valid inference-to the adaptive IV setting with noncompliance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a57922c-6efd-457c-9754-cd881ec194a4Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Statistical Inference with M-Estimators on Adaptively Collected DataKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2021 · 66 citations
- The Adaptive Doubly Robust Estimator and a Paradox Concerning Logging PolicyMasahiro Kato, Kenichiro McAlinn, Shota YasuiNeurIPS 2021 · 23 citations
- CLIP-OGD: An Experimental Design for Adaptive Neyman Allocation in Sequential ExperimentsJessica Dai, Paula Gradu, Christopher HarshawNeurIPS 2023 · 21 citations
- Efficient Online Estimation of Causal Effects by Deciding What to ObserveShantanu Gupta, Zachary C. Lipton, David ChildersNeurIPS 2021 · 19 citations
- Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate ChoiceMasahiro Kato, Akihiro Oga, Wataru Komatsubara, Ryo InokuchiICML 2024 · 12 citations
Related papers
- Adaptive Instrument Design for Indirect ExperimentsYash Chandak, Shiv Shankar, Vasilis Syrgkanis, Emma BrunskillICLR 2024 · 5 citations
- Beyond the Average: Distributional Causal Inference under Imperfect ComplianceUndral Byambadalai, Tomu Hirata, Tatsushi Oka, Shota YasuiNeurIPS 2025 · 3 citations
- A Non-parametric Direct Learning Approach to Heterogeneous Treatment Effect Estimation under Unmeasured ConfoundingXinhai Zhang, Xingye QiaoNeurIPS 2024 · 1 citation
- Off-policy estimation with adaptively collected data: the power of online learningJeonghwan Lee, Cong MaNeurIPS 2024 · 4 citations
- Fair Adaptive ExperimentsWaverly Wei, Xinwei Ma, Jingshen WangNeurIPS 2023 · 7 citations
