Unifying Offline Causal Inference and Online Bandit Learning for Data Driven Decision
Ye Li, Hong Xie, Yishi Lin, John C. S. Lui
Abstract
A fundamental question for companies with large amount of logged data is: How to use such logged data together with incoming streaming data to make good decisions? Many companies currently make decisions via online A/B tests, but wrong decisions during testing hurt users’ experiences and cause irreversible damage. A typical alternative is offline causal inference, which analyzes logged data alone to make decisions. However, these decisions are not adaptive to the new incoming data, and so a wrong decision will continuously hurt users’ experiences. To overcome the aforementioned limitations, we propose a framework to unify offline causal inference algorithms (e.g., weighting, matching) and online learning algorithms (e.g., UCB, LinUCB). We propose novel algorithms and derive bounds on the decision accuracy via the notion of “regret”. We derive the first upper regret bound for forest-based online bandit algorithms. Experiments on two real datasets show that our algorithms outperform other algorithms that use only logged data or online feedbacks, or algorithms that do not use the data properly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db61ca55-d21a-481b-9a7e-ac2f5e206e31Cited by top-tier papers4
- Learning to price with resource constraints: from full information to machine-learned pricesRuicheng Ao, Jiashuo Jiang, David Simchi-LeviNeurIPS 2025 · 4 citations
- Contextual Online Pricing with (Biased) Offline DataYixuan Zhang, Ruihao Zhu, Qiaomin XieNeurIPS 2025 · 3 citations
- Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal ApproachWen Huang, Xintao WuAAAI 2024 · 2 citations
- Conformal Counterfactual Inference under Hidden ConfoundingZonghao Chen, Ruocheng Guo, Jean-Francois Ton, Yang LiuKDD 2024 · 2 citations
Related papers
- Online Multi-Armed Bandits with Adaptive InferenceMaria Dimakopoulou, Zhimei Ren, Zhengyuan ZhouNeurIPS 2021 · 47 citations
- Cascading Bandits: Optimizing Recommendation Frequency in Delayed Feedback EnvironmentsDairui Wang, Junyu Cao, Yan Zhang, Wei QiNeurIPS 2023 · 2 citations
- Statistical Inference on Multi-armed Bandits with Delayed FeedbackLei Shi, Jingshen Wang, Tianhao WuICML 2023 · 7 citations
- Achieving Counterfactual Fairness for Causal BanditWen Huang, Lu Zhang, Xintao WuAAAI 2022 · 33 citations
- Pessimistic Data Integration for Policy EvaluationXiangkun Wu, Ting Li, Gholamali Aminian, Armin Behnamnia et al.NeurIPS 2025 · 2 citations
