Markovian Interference in Experiments
Vivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew Zheng
摘要
We consider experiments in dynamical systems where interventions on some experimental units impact other units through a limiting constraint (such as a limited inventory). Despite outsize practical importance, the best estimators for this `Markovian' interference problem are largely heuristic in nature, and their bias is not well understood. We formalize the problem of inference in such experiments as one of policy evaluation. Off-policy estimators, while unbiased, apparently incur a large penalty in variance relative to state-of-the-art heuristics. We introduce an on-policy estimator: the Differences-In-Q's (DQ) estimator. We show that the DQ estimator can in general have exponentially smaller variance than off-policy evaluation. At the same time, its bias is second order in the impact of the intervention. This yields a striking bias-variance tradeoff so that the DQ estimator effectively dominates state-of-the-art alternatives. From a theoretical perspective, we introduce three separate novel techniques that are of independent interest in the theory of Reinforcement Learning (RL). Our empirical evaluation includes a set of experiments on a city-scale ride-hailing simulator.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou 等NeurIPS 2023 · 被引用 21 次
- Higher-Order Causal Message Passing for Experimentation with Complex InterferenceMohsen Bayati, Yuwei Luo, William Overman, Mohamad Sadegh Shirani Faradonbeh 等NeurIPS 2024 · 被引用 9 次
- Reducing Symbiosis Bias through Better A/B Tests of Recommendation AlgorithmsJennifer Brennan, Yahu Cong, Yiwei Yu, Lina Lin 等WWW 2025 · 被引用 8 次
- On the Statistical Benefits of Temporal Difference LearningDavid Cheikhi, Daniel RussoICML 2023 · 被引用 6 次
- Pricing Experimental Design: Causal Effect, Expected Revenue and Tail RiskDavid Simchi-Levi, Chonghuan WangICML 2023 · 被引用 6 次
它引用的顶会 Paper7
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 被引用 199 次
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 125 次
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 被引用 82 次
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou 等ICLR 2020 · 被引用 72 次
- Adaptive Experimental Design with Temporal Interference: A Maximum Likelihood ApproachPeter W. Glynn, Ramesh Johari, Mohammad RasouliNeurIPS 2020 · 被引用 45 次
相关 Paper
- Controlling Underestimation Bias in Reinforcement Learning via Quasi-median OperationWei Wei, Yujia Zhang, Jiye Liang, Lin Li 等AAAI 2022 · 被引用 20 次
- Reward Shaping Control Variates for Off-Policy Evaluation Under Sparse RewardsRitam Majumdar, Finale Doshi-Velez, Sonali ParbhooICML 2026 · 被引用 3 次
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient EstimatorRamesh Johari, Tianyi Peng, Wenqian XingICML 2026
- Experimentation for Different Scheduling Policies on Queues: Mixed Differences-in-Q Estimators Based on Little's LawNanshan Jia, Ramesh Johari, Nian Si, Zeyu ZhengKDD 2026
