Online Limited Memory Neural-Linear Bandits with Likelihood Matching
Ofir Nabati, Tom Zahavy, Shie Mannor
摘要
We study neural-linear bandits for solving problems where both exploration and representation learning play an important role. Neural-linear bandits harnesses the representation power of Deep Neural Networks (DNNs) and combines it with efficient exploration mechanisms by leveraging uncertainty estimation of the model, designed for linear contextual bandits on top of the last hidden layer. In order to mitigate the problem of representation change during the process, new uncertainty estimations are computed using stored data from an unlimited buffer. Nevertheless, when the amount of stored data is limited, a phenomenon called catastrophic forgetting emerges. To alleviate this, we propose a likelihood matching algorithm that is resilient to catastrophic forgetting and is completely online. We applied our algorithm, Limited Memory Neural-Linear with Likelihood Matching (NeuralLinear-LiM2) on a variety of datasets and observed that our algorithm achieves comparable performance to the unlimited memory approach while exhibits resilience to catastrophic forgetting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Sample-Then-Optimize Batch Neural Thompson SamplingZhongxiang Dai, Yao Shu, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2022 · 被引用 33 次
- Anti-Concentrated Confidence Bonuses For Scalable ExplorationJordan T. Ash, Cyril Zhang, Surbhi Goel, Akshay Krishnamurthy 等ICLR 2022 · 被引用 9 次
- Meta Clustering of Neural BanditsYikun Ban, Yunzhe Qi, Tianxin Wei, Lihui Liu 等KDD 2024 · 被引用 6 次
- Representation-Driven Reinforcement LearningOfir Nabati, Guy Tennenholtz, Shie MannorICML 2023 · 被引用 3 次
- Federated Neural BanditsZhongxiang Dai, Yao Shu, Arun Verma, Flint Xiaofeng Fan 等ICLR 2023 · 被引用 2 次
它引用的顶会 Paper3
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 被引用 329 次
- Neural Thompson SamplingWeitong Zhang, Dongruo Zhou, Lihong Li, Quanquan GuICLR 2021 · 被引用 152 次
- Neural Contextual Bandits with Deep Representation and Shallow ExplorationPan Xu, Zheng Wen, Handong Zhao, Quanquan GuICLR 2022 · 被引用 90 次
相关 Paper
- Learning Neural Contextual Bandits through Perturbed RewardsYiling Jia, Weitong Zhang, Dongruo Zhou, Quanquan Gu 等ICLR 2022 · 被引用 20 次
- Online Learning in Contextual Bandits using Gated Linear NetworksEren Sezener, Marcus Hutter, David Budden, Jianan Wang 等NeurIPS 2020 · 被引用 10 次
- Scalable Representation Learning in Linear Contextual Bandits with Constant Regret GuaranteesAndrea Tirinzoni, Matteo Papini, Ahmed Touati, Alessandro Lazaric 等NeurIPS 2022 · 被引用 7 次
- Learning without Prejudices: Continual Unbiased Learning via Benign and Malignant ForgettingMyeongho Jeon, Hyoje Lee, Yedarm Seong, Myungjoo KangICLR 2023
- Class-Incremental Learning by Knowledge Distillation with Adaptive Feature ConsolidationMinsoo Kang, Jaeyoo Park, Bohyung HanCVPR 2022 · 被引用 189 次
