Online Limited Memory Neural-Linear Bandits with Likelihood Matching
Ofir Nabati, Tom Zahavy, Shie Mannor
Abstract
We study neural-linear bandits for solving problems where both exploration and representation learning play an important role. Neural-linear bandits harnesses the representation power of Deep Neural Networks (DNNs) and combines it with efficient exploration mechanisms by leveraging uncertainty estimation of the model, designed for linear contextual bandits on top of the last hidden layer. In order to mitigate the problem of representation change during the process, new uncertainty estimations are computed using stored data from an unlimited buffer. Nevertheless, when the amount of stored data is limited, a phenomenon called catastrophic forgetting emerges. To alleviate this, we propose a likelihood matching algorithm that is resilient to catastrophic forgetting and is completely online. We applied our algorithm, Limited Memory Neural-Linear with Likelihood Matching (NeuralLinear-LiM2) on a variety of datasets and observed that our algorithm achieves comparable performance to the unlimited memory approach while exhibits resilience to catastrophic forgetting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Sample-Then-Optimize Batch Neural Thompson SamplingZhongxiang Dai, Yao Shu, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2022 · 33 citations
- Anti-Concentrated Confidence Bonuses For Scalable ExplorationJordan T. Ash, Cyril Zhang, Surbhi Goel, Akshay Krishnamurthy et al.ICLR 2022 · 9 citations
- Meta Clustering of Neural BanditsYikun Ban, Yunzhe Qi, Tianxin Wei, Lihui Liu et al.KDD 2024 · 6 citations
- Representation-Driven Reinforcement LearningOfir Nabati, Guy Tennenholtz, Shie MannorICML 2023 · 3 citations
- Federated Neural BanditsZhongxiang Dai, Yao Shu, Arun Verma, Flint Xiaofeng Fan et al.ICLR 2023 · 2 citations
Builds on3
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 329 citations
- Neural Thompson SamplingWeitong Zhang, Dongruo Zhou, Lihong Li, Quanquan GuICLR 2021 · 152 citations
- Neural Contextual Bandits with Deep Representation and Shallow ExplorationPan Xu, Zheng Wen, Handong Zhao, Quanquan GuICLR 2022 · 90 citations
Related papers
- Learning Neural Contextual Bandits through Perturbed RewardsYiling Jia, Weitong Zhang, Dongruo Zhou, Quanquan Gu et al.ICLR 2022 · 20 citations
- Online Learning in Contextual Bandits using Gated Linear NetworksEren Sezener, Marcus Hutter, David Budden, Jianan Wang et al.NeurIPS 2020 · 10 citations
- Scalable Representation Learning in Linear Contextual Bandits with Constant Regret GuaranteesAndrea Tirinzoni, Matteo Papini, Ahmed Touati, Alessandro Lazaric et al.NeurIPS 2022 · 7 citations
- Learning without Prejudices: Continual Unbiased Learning via Benign and Malignant ForgettingMyeongho Jeon, Hyoje Lee, Yedarm Seong, Myungjoo KangICLR 2023
- Class-Incremental Learning by Knowledge Distillation with Adaptive Feature ConsolidationMinsoo Kang, Jaeyoo Park, Bohyung HanCVPR 2022 · 189 citations
