Neural Contextual Bandits with Deep Representation and Shallow Exploration
Pan Xu, Zheng Wen, Handong Zhao, Quanquan Gu
Abstract
We study a general class of contextual bandits, where each context-action pair is associated with a raw feature vector, but the reward generating function is unknown. We propose a novel learning algorithm that transforms the raw feature vector using the last hidden layer of a deep ReLU neural network (deep representation learning), and uses an upper confidence bound (UCB) approach to explore in the last linear layer (shallow exploration). We prove that under standard assumptions, our proposed algorithm achieves finite-time regret, where is the learning time horizon. Compared with existing neural contextual bandit algorithms, our approach is computationally much more efficient since it only needs to explore in the last layer of the deep neural network.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6d5d15f-e59a-46f0-b846-42a00fc63dc1Cited by top-tier papers36
- EE-Net: Exploitation-Exploration Neural Networks in Contextual BanditsYikun Ban, Yuchen Yan, Arindam Banerjee, Jingrui HeICLR 2022 · 62 citations
- Offline Neural Contextual Bandits: Pessimism, Optimization and GeneralizationThanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, Svetha VenkateshICLR 2022 · 35 citations
- Contextual Bandits with Large Action Spaces: Made PracticalYinglun Zhu, Dylan J. Foster, John Langford, Paul MineiroICML 2022 · 34 citations
- Langevin Monte Carlo for Contextual BanditsPan Xu, Hongkai Zheng, Eric V. Mazumdar, Kamyar Azizzadenesheli et al.ICML 2022 · 34 citations
- Sample-Then-Optimize Batch Neural Thompson SamplingZhongxiang Dai, Yao Shu, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2022 · 33 citations
Builds on4
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 329 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- Neural Thompson SamplingWeitong Zhang, Dongruo Zhou, Lihong Li, Quanquan GuICLR 2021 · 152 citations
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 79 citations
Related papers
- Learning Neural Contextual Bandits through Perturbed RewardsYiling Jia, Weitong Zhang, Dongruo Zhou, Quanquan Gu et al.ICLR 2022 · 20 citations
- Neural Combinatorial Clustered Bandits for Recommendation SystemsBaran Atalar, Carlee Joe-WongAAAI 2025 · 4 citations
- Stochastic Bandits with ReLU Neural NetworksKan Xu, Hamsa Bastani, Surbhi Goel, Osbert BastaniICML 2024 · 1 citation
- Provable General Function Class Representation Learning in Multitask Bandits and MDPRui Lu, Andrew Zhao, Simon S. Du, Gao HuangNeurIPS 2022 · 11 citations
- Scalable Representation Learning in Linear Contextual Bandits with Constant Regret GuaranteesAndrea Tirinzoni, Matteo Papini, Ahmed Touati, Alessandro Lazaric et al.NeurIPS 2022 · 7 citations
