Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal Approach
Wen Huang, Xintao Wu
Abstract
This paper studies bandit problems where an agent has access to offline data that might be utilized to potentially improve the estimation of each arm’s reward distribution. A major obstacle in this setting is the existence of compound biases from the observational data. Ignoring these biases and blindly fitting a model with the biased data could even negatively affect the online learning phase. In this work, we formulate this problem from a causal perspective. First, we categorize the biases into confounding bias and selection bias based on the causal structure they imply. Next, we extract the causal bound for each arm that is robust towards compound biases from biased observational data. The derived bounds contain the ground truth mean reward and can effectively guide the bandit agent to learn a nearly-optimal decision policy. We also conduct regret analysis in both contextual and non-contextual bandit settings and show that prior causal bounds could help consistently reduce the asymptotic regret.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 5 citations
- Non-Stationary Structural Causal BanditsYeahoon Kwon, Yesong Choe, Soungmin Park, Neil Dhir et al.NeurIPS 2025
Builds on4
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 329 citations
- Bounding Causal Effects on Continuous OutcomeJunzhe Zhang, Elias BareinboimAAAI 2021 · 49 citations
- Stochastic bandits for multi-platform budget optimization in online advertisingVashist Avadhanula, Riccardo Colini-Baldeschi, Stefano Leonardi, Karthik Abinav Sankararaman et al.WWW 2021 · 43 citations
- Unifying Offline Causal Inference and Online Bandit Learning for Data Driven DecisionYe Li, Hong Xie, Yishi Lin, John C. S. LuiWWW 2021 · 17 citations
Related papers
- Counterfactual Structural Causal BanditsMin Woo Park, Sanghack LeeICLR 2026 · 1 citation
- Transportability for Bandits with Data from Different EnvironmentsAlexis Bellot, Alan Malek, Silvia ChiappaNeurIPS 2023 · 11 citations
- Automatic Reward Shaping from Confounded Offline DataMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2025
- Causal Bandits: The Pareto Optimal Frontier of Adaptivity, a Reduction to Linear Bandits, and Limitations around Unknown MarginalsZiyi Liu, Idan Attias, Daniel M. RoyICML 2024 · 2 citations
- Leveraging (Biased) Information: Multi-armed Bandits with Offline DataWang Chi Cheung, Lixing LyuICML 2024 · 3 citations
