Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator
Ramesh Johari, Tianyi Peng, Wenqian Xing
Abstract
Randomized experiments (or A/B tests) are widely used to evaluate interventions in dynamic systems such as recommendation platforms, marketplaces, and digital health. In these settings, interventions affect both current and future system states, so estimating the global average treatment effect (GATE) requires accounting for temporal dynamics, which is especially challenging in the presence of nonstationarity; existing approaches suffer from high bias, high variance, or both. In this paper, we address this challenge via the novel Truncated Policy Gradient (TPG) estimator, which replaces instantaneous outcomes with short-horizon outcome trajectories. The estimator admits a policy gradient interpretation: it is a truncation of the first-order approximation to the GATE, yielding provable reductions in bias and variance in nonstationary Markovian settings. We further establish a central limit theorem for the TPG estimator and develop a consistent variance estimator that remains valid under nonstationarity with single-trajectory data. We validate our theory with two real-world case studies. The results show that relative to existing approaches, a well-calibrated TPG estimator can achieve a favorable balance between bias and variance in nonstationary settings, highlighting the value of the policy-gradient perspective for designing effective estimators under complex dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) OptimismWang Chi Cheung, David Simchi-Levi, Ruihao ZhuICML 2020 · 114 citations
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 55 citations
- Markovian Interference in ExperimentsVivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew ZhengNeurIPS 2022 · 52 citations
- Adaptive Experimental Design with Temporal Interference: A Maximum Likelihood ApproachPeter W. Glynn, Ramesh Johari, Mohammad RasouliNeurIPS 2020 · 45 citations
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov et al.NeurIPS 2023 · 31 citations
Related papers
- Modeling Covariate Transition for Efficient Estimation of Longitudinal Treatment Effects in Randomized ExperimentsNaoki Chihara, Tatsushi Oka, Yasuko Matsubara, Yasushi Sakurai et al.ICML 2026
- Non-stationary Experimental Design under Linear TrendsDavid Simchi-Levi, Chonghuan Wang, Zeyu ZhengNeurIPS 2023 · 6 citations
- Optimized Covariance Design for AB Test on Social Network under InterferenceQianyi Chen, Bo Li, Lu Deng, Yong WangNeurIPS 2023 · 6 citations
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Non-stationary A/B TestsYuhang Wu, Zeyu Zheng, Guangyu Zhang, Zuohua Zhang et al.KDD 2022 · 5 citations
