Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning
Xiangkun Wu, Qianglin Wen, Yingying Zhang, Hongtu Zhu, Ting Li, Chengchun Shi
Abstract
A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where treatments are sequentially assigned over time, remains challenging. Existing designs suffer from two limitations: (i) they do not fully leverage the entire history for treatment allocation; (ii) they rely on strong assumptions to approximate the objective function (e.g., the mean squared error of the estimated treatment effect) for optimizing the design. We first establish an impossibility theorem showing that failure to condition on the full history leads to suboptimal designs, due to the dynamic dependencies in time series experiments. To address both limitations simultaneously, we next propose a transformer reinforcement learning (RL) approach which leverages transformers to condition treatment allocation on the entire history and employs RL to directly optimize the MSE without relying on restrictive assumptions. Empirical evaluations on synthetic data, a publicly available dispatch simulator, and a real-world ridesharing dataset demonstrate that our proposal consistently outperforms existing designs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9dcbe9d5-2a22-4e55-9662-063c6d231d68Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 3,619 citations
- Causal Discovery with Reinforcement LearningShengyu Zhu, Ignavier Ng, Zhitang ChenICLR 2020 · 285 citations
- Deep Adaptive Design: Amortizing Sequential Bayesian Experimental DesignAdam Foster, Desi R. Ivanova, Ilyas Malik, Tom RainforthICML 2021 · 119 citations
- Optimizing Sequential Experimental Design with Deep Reinforcement LearningTom Blau, Edwin V. Bonilla, Iadine Chades, Amir DezfouliICML 2022 · 62 citations
- Markovian Interference in ExperimentsVivek F. Farias, Andrew A. Li, Tianyi Peng, Andrew ZhengNeurIPS 2022 · 52 citations
Related papers
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- DRIVE: Distributional and Retrieval-Augmented Bidding with Value EvaluationMiduo Cui, Haochen Wang, Shangqin Mao, Xun Yang et al.ICML 2026 · 1 citation
- Emergent Agentic Transformer from Chain of Hindsight ExperienceHao Liu, Pieter AbbeelICML 2023 · 35 citations
- TransformerLight: A Novel Sequence Modeling Based Traffic Signaling Mechanism via Gated TransformerQiang Wu, Mingyuan Li, Jun Shen, Linyuan Lü et al.KDD 2023 · 17 citations
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 77 citations
