Learning Across the Gap: Hybrid Multi-armed Bandits with Heterogeneous Offline and Online Data
Qijia He, Minghan Wang, Xutong Liu, Zhiyong Wang, Fang Kong
Abstract
The multi-armed bandit (MAB) is a fundamental online decision-making framework that has been extensively studied over the past two decades. To mitigate the high cost and slow convergence of purely online learning, modern MAB approaches have explored hybrid paradigms that leverage offline data to warm-start online learning. However, existing approaches face a significant limitation by assuming that the offline and online data are homogeneous-they share the same feedback structure and are drawn from the same underlying distribution. This assumption is often violated in practice, where offline data often originate from diverse sources and evolving environments, resulting in feedback heterogeneity and distributional shifts. In this work, we tackle the challenge of learning across this offline-online gap by developing a general hybrid bandit framework that incorporates heterogeneous offline data to improve online performance. We study two hybrid settings: (1) using reward-based offline data to accelerate online learning in preference-based bandits (i.e., dueling bandits), and (2) using preference-based offline data to improve online standard MAB algorithms. For both settings, we design novel algorithms and derive tight regret bounds that match or improve upon existing benchmarks despite heterogeneity. Empirical evaluations on both synthetic and real-world datasets further show the superior performance of our proposed methods over baseline algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70ffd898-4c1d-4f68-9c3e-20fdfb67a1fcBuilds on9
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Conversational Contextual Bandit: Algorithm and ApplicationXiaoying Zhang, Hong Xie, Hang Li, John C. S. LuiWWW 2020 · 97 citations
- Online Pricing with Offline Data: Phase Transition and Inverse Square LawJinzhi Bu, David Simchi-Levi, Yunzong XuICML 2020 · 40 citations
- Adversarial Dueling BanditsAadirupa Saha, Tomer Koren, Yishay MansourICML 2021 · 35 citations
- Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative PreferencesAadirupa Saha, Pierre GaillardICML 2022 · 30 citations
Related papers
- Leveraging (Biased) Information: Multi-armed Bandits with Offline DataWang Chi Cheung, Lixing LyuICML 2024 · 3 citations
- Robustly Improving Bandit Algorithms with Confounded and Selection Biased Offline Data: A Causal ApproachWen Huang, Xintao WuAAAI 2024 · 2 citations
- Adaptive Policy Learning for Offline-to-Online Reinforcement LearningHan Zheng, Xufang Luo, Pengfei Wei, Xuan Song et al.AAAI 2023 · 47 citations
- Online Clustering of Dueling BanditsZhiyong Wang, Jiahang Sun, Mingze Kong, Jize Xie et al.ICML 2025
- Learning Preferences without Interaction for Cooperative AI: A Hybrid Offline-Online ApproachHaitong Ma, Haoran Yu, Haobo Fu, Shuai LiNeurIPS 2025
