Finite-Time Analysis of Whittle Index based Q-Learning for Restless Multi-Armed Bandits with Neural Network Function Approximation
Guojun Xiong, Jian Li
摘要
Whittle index policy is a heuristic to the intractable restless multi-armed bandits (RMAB) problem. Although it is provably asymptotically optimal, finding Whittle indices remains difficult. In this paper, we present Neural-Q-Whittle, a Whittle index based Q-learning algorithm for RMAB with neural network function approximation, which is an example of nonlinear two-timescale stochastic approximation with Q-function values updated on a faster timescale and Whittle indices on a slower timescale. Despite the empirical success of deep Q-learning, the non-asymptotic convergence rate of Neural-Q-Whittle, which couples neural networks with two-timescale Q-learning largely remains unclear. This paper provides a finite-time analysis of Neural-Q-Whittle, where data are generated from a Markov chain, and Q-function is approximated by a ReLU neural network. Our analysis leverages a Lyapunov drift approach to capture the evolution of two coupled parameters, and the nonlinearity in value function approximation further requires us to characterize the approximation error. Combing these provide Neural-Q-Whittle with O(1/k 2/3 ) convergence rate, where k is the number of iterations. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Online Restless Multi-Armed Bandits with Long-Term Fairness ConstraintsShufan Wang, Guojun Xiong, Jian LiAAAI 2024 · 被引用 11 次
- The Bandit Whisperer: Communication Learning for Restless BanditsYunfan Zhao, Tonghan Wang, Dheeraj Mysore Nagaraj, Aparna Taneja 等AAAI 2025 · 被引用 6 次
- Projection-based Lyapunov method for fully heterogeneous weakly-coupled MDPsXiangcheng Zhang, Yige Hong, Weina WangNeurIPS 2025 · 被引用 2 次
- GINO-Q: Learning an Asymptotically Optimal Index Policy for Restless Multi-armed BanditsGongpu Chen, Soung Chang Liew, Deniz GündüzAAAI 2026 · 被引用 1 次
- Achieving 𝒪(1/N) Optimality Gap in Restless Bandits through Gaussian ApproximationChen Yan, Weina Wang, Lei YingNeurIPS 2025
它引用的顶会 Paper14
- Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision ProcessesChen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Hiteshi Sharma 等ICML 2020 · 被引用 120 次
- Collapsing Bandits and Their Application to Public Health InterventionAditya Mate, Jackson A. Killian, Haifeng Xu, Andrew Perrault 等NeurIPS 2020 · 被引用 83 次
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 被引用 82 次
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 被引用 79 次
- A Tale of Two-Timescale Reinforcement Learning with the Tightest Finite-Time BoundGal Dalal, Balázs Szörényi, Gugan ThoppeAAAI 2020 · 被引用 59 次
相关 Paper
- NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RLKhaled Nakhleh, Santosh Ganji, Ping-Chun Hsieh, I-Hong Hou 等NeurIPS 2021 · 被引用 52 次
- Whittle Index with Multiple Actions and State Constraint for Inventory ManagementChuheng Zhang, Xiangsen Wang, Wei Jiang, Xianliang Yang 等ICLR 2024 · 被引用 10 次
- Q-Learning Lagrange Policies for Multi-Action Restless BanditsJackson A. Killian, Arpita Biswas, Sanket Shah, Milind TambeKDD 2021 · 被引用 12 次
- DeepTOP: Deep Threshold-Optimal Policy for MDPs and RMABsKhaled Nakhleh, I-Hong HouNeurIPS 2022 · 被引用 12 次
- Optimistic Whittle Index Policy: Online Learning for Restless BanditsKai Wang, Lily Xu, Aparna Taneja, Milind TambeAAAI 2023 · 被引用 31 次
