Leveraging Demonstrations to Improve Online Learning: Quality Matters
Botao Hao, Rahul Jain, Tor Lattimore, Benjamin Van Roy, Zheng Wen
摘要
We investigate the extent to which offline demonstration data can improve online learning. It is natural to expect some improvement, but the question is how, and by how much? We show that the degree of improvement must depend on the quality of the demonstration data. To generate portable insights, we focus on Thompson sampling (TS) applied to a multi-armed bandit as a prototypical online learning algorithm and model. The demonstration data is generated by an expert with a given competence level, a notion we introduce. We propose an informed TS algorithm that utilizes the demonstration data in a coherent way through Bayes' rule and derive a prior-dependent Bayesian regret bound. This offers insight into how pretraining can greatly improve online performance and how the degree of improvement increases with the expert's competence level. We also develop a practical, approximate informed TS algorithm through Bayesian bootstrapping and show substantial empirical regret reduction through experiments. * Equal contribution 1 Deepmind 2 University of Southern California, work done by Rahul Jain while at DeepMind.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Across the Gap: Hybrid Multi-armed Bandits with Heterogeneous Offline and Online DataQijia He, Minghan Wang, Xutong Liu, Zhiyong Wang 等NeurIPS 2025 · 被引用 6 次
- Leveraging (Biased) Information: Multi-armed Bandits with Offline DataWang Chi Cheung, Lixing LyuICML 2024 · 被引用 3 次
- Sequential Decision Making with Expert Demonstrations under Unobserved HeterogeneityVahid Balazadeh Meresht, Keertana Chidambaram, Viet Nguyen, Rahul G. Krishnan 等NeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong 等NeurIPS 2021 · 被引用 207 次
- Epistemic Neural NetworksIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla 等NeurIPS 2023 · 被引用 142 次
- Meta-Thompson SamplingBranislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih-Wei Hsu 等ICML 2021 · 被引用 74 次
相关 Paper
- Thompson Sampling for Robust Transfer in Multi-Task BanditsZhi Wang, Chicheng Zhang, Kamalika ChaudhuriICML 2022 · 被引用 7 次
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak 等NeurIPS 2023 · 被引用 170 次
- Contextual Thompson Sampling via Generation of Missing DataKelly W. Zhang, Tiffany Tianhui Cai, Hongseok Namkoong, Daniel RussoNeurIPS 2025 · 被引用 5 次
- No Regrets for Learning the Prior in BanditsSoumya Basu, Branislav Kveton, Manzil Zaheer, Csaba SzepesváriNeurIPS 2021 · 被引用 39 次
- GuideBoot: Guided Bootstrap for Deep Contextual Banditsin Online AdvertisingFeiyang Pan, Haoming Li, Xiang Ao, Wei Wang 等WWW 2021
