Markovian Interference in Experiments
Vivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew Zheng
Abstract
We consider experiments in dynamical systems where interventions on some experimental units impact other units through a limiting constraint (such as a limited inventory). Despite outsize practical importance, the best estimators for this `Markovian' interference problem are largely heuristic in nature, and their bias is not well understood. We formalize the problem of inference in such experiments as one of policy evaluation. Off-policy estimators, while unbiased, apparently incur a large penalty in variance relative to state-of-the-art heuristics. We introduce an on-policy estimator: the Differences-In-Q's (DQ) estimator. We show that the DQ estimator can in general have exponentially smaller variance than off-policy evaluation. At the same time, its bias is second order in the impact of the intervention. This yields a striking bias-variance tradeoff so that the DQ estimator effectively dominates state-of-the-art alternatives. From a theoretical perspective, we introduce three separate novel techniques that are of independent interest in the theory of Reinforcement Learning (RL). Our empirical evaluation includes a set of experiments on a city-scale ride-hailing simulator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Higher-Order Causal Message Passing for Experimentation with Complex InterferenceMohsen Bayati, Yuwei Luo, William Overman, Mohamad Sadegh Shirani Faradonbeh et al.NeurIPS 2024 · 9 citations
- Reducing Symbiosis Bias through Better A/B Tests of Recommendation AlgorithmsJennifer Brennan, Yahu Cong, Yiwei Yu, Lina Lin et al.WWW 2025 · 8 citations
- On the Statistical Benefits of Temporal Difference LearningDavid Cheikhi, Daniel RussoICML 2023 · 6 citations
- Pricing Experimental Design: Causal Effect, Expected Revenue and Tail RiskDavid Simchi-Levi, Chonghuan WangICML 2023 · 6 citations
Builds on7
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li et al.NeurIPS 2020 · 125 citations
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou et al.ICLR 2020 · 72 citations
- Adaptive Experimental Design with Temporal Interference: A Maximum Likelihood ApproachPeter W. Glynn, Ramesh Johari, Mohammad RasouliNeurIPS 2020 · 45 citations
Related papers
- Controlling Underestimation Bias in Reinforcement Learning via Quasi-median OperationWei Wei, Yujia Zhang, Jiye Liang, Lin Li et al.AAAI 2022 · 20 citations
- Reward Shaping Control Variates for Off-Policy Evaluation Under Sparse RewardsRitam Majumdar, Finale Doshi-Velez, Sonali ParbhooICML 2026 · 3 citations
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient EstimatorRamesh Johari, Tianyi Peng, Wenqian XingICML 2026
- Experimentation for Different Scheduling Policies on Queues: Mixed Differences-in-Q Estimators Based on Little's LawNanshan Jia, Ramesh Johari, Nian Si, Zeyu ZhengKDD 2026
