Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal Approach
Wen Huang, Xintao Wu
摘要
This paper studies bandit problems where an agent has access to offline data that might be utilized to potentially improve the estimation of each arm’s reward distribution. A major obstacle in this setting is the existence of compound biases from the observational data. Ignoring these biases and blindly fitting a model with the biased data could even negatively affect the online learning phase. In this work, we formulate this problem from a causal perspective. First, we categorize the biases into confounding bias and selection bias based on the causal structure they imply. Next, we extract the causal bound for each arm that is robust towards compound biases from biased observational data. The derived bounds contain the ground truth mean reward and can effectively guide the bandit agent to learn a nearly-optimal decision policy. We also conduct regret analysis in both contextual and non-contextual bandit settings and show that prior causal bounds could help consistently reduce the asymptotic regret.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 被引用 5 次
- Non-Stationary Structural Causal BanditsYeahoon Kwon, Yesong Choe, Soungmin Park, Neil Dhir 等NeurIPS 2025
它引用的顶会 Paper4
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 被引用 329 次
- Bounding Causal Effects on Continuous OutcomeJunzhe Zhang, Elias BareinboimAAAI 2021 · 被引用 49 次
- Stochastic bandits for multi-platform budget optimization in online advertisingVashist Avadhanula, Riccardo Colini-Baldeschi, Stefano Leonardi, Karthik Abinav Sankararaman 等WWW 2021 · 被引用 43 次
- Unifying Offline Causal Inference and Online Bandit Learning for Data Driven DecisionYe Li, Hong Xie, Yishi Lin, John C. S. LuiWWW 2021 · 被引用 17 次
相关 Paper
- Counterfactual Structural Causal BanditsMin Woo Park, Sanghack LeeICLR 2026 · 被引用 1 次
- Transportability for Bandits with Data from Different EnvironmentsAlexis Bellot, Alan Malek, Silvia ChiappaNeurIPS 2023 · 被引用 11 次
- Automatic Reward Shaping from Confounded Offline DataMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2025
- Causal Bandits: The Pareto Optimal Frontier of Adaptivity, a Reduction to Linear Bandits, and Limitations around Unknown MarginalsZiyi Liu, Idan Attias, Daniel M. RoyICML 2024 · 被引用 2 次
- Leveraging (Biased) Information: Multi-armed Bandits with Offline DataWang Chi Cheung, Lixing LyuICML 2024 · 被引用 3 次
