Reinforcement learning with combinatorial actions for coupled restless bandits
Lily Xu, Bryan Wilder, Elias Boutros Khalil, Milind Tambe
Abstract
Reinforcement learning (RL) has increasingly been applied to solve real-world planning problems, with progress in handling large state spaces and time horizons. However, a key bottleneck in many domains is that RL methods cannot accommodate large, combinatorially structured action spaces. In such settings, even representing the set of feasible actions at a single step may require a complex discrete optimization formulation. We leverage recent advances in embedding trained neural networks into optimization problems to propose SEQUOIA, an RL algorithm that directly optimizes for long-term reward over the feasible action space. Our approach embeds a Q-network into a mixed-integer program to select a combinatorial action in each timestep. Here, we focus on planning over restless bandits, a class of planning problems which capture many real-world examples of sequential decision making. We introduce CORMAB, a broader class of restless bandits with combinatorial actions that cannot be decoupled across the arms of the restless bandit, requiring direct solving over the joint, exponentially large action space. We empirically validate SEQUOIA on four novel restless bandit problems with combinatorial constraints: multiple interventions, path constraints, bipartite matching, and capacity constraints. Our approach significantly outperforms existing methods-which cannot address sequential planning and combinatorial selection simultaneously-by an average of 24.8% on these difficult instances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8535cfde-d857-4442-ae04-3abb8d34d8b6Cited by top-tier papers4
- Structured Reinforcement Learning for Combinatorial Decision-MakingHeiko Hoppe, Léo Baty, Louis Bouvier, Axel Parmentier et al.NeurIPS 2025 · 12 citations
- Global Rewards in Restless Multi-Armed BanditsNaveen Raman, Zheyuan Shi, Fei FangNeurIPS 2024 · 10 citations
- Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial ActionsLingkai Kong, Anagha Satish, Hezi Jiang, Akseli Kangaslahti et al.ICML 2026 · 1 citation
- Combinatorial Rising BanditsSeockbean Song, Youngsik Yoon, Siwei Wang, Wei Chen et al.ICLR 2026
Builds on11
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Reinforcement Learning for Integer Programming: Learning to CutYunhao Tang, Shipra Agrawal, Yuri FaenzaICML 2020 · 224 citations
- Exploratory Combinatorial Optimization with Reinforcement LearningThomas D. Barrett, William R. Clements, Jakob N. Foerster, A. I. LvovskyAAAI 2020 · 218 citations
- Reinforcement Learning with Combinatorial Actions: An Application to Vehicle RoutingArthur Delarue, Ross Anderson, Christian TjandraatmadjaNeurIPS 2020 · 127 citations
- Winner Takes It All: Training Performant RL Populations for Combinatorial OptimizationNathan Grinsztajn, Daniel Furelos-Blanco, Shikha Surana, Clément Bonnet et al.NeurIPS 2023 · 79 citations
Related papers
- CAQL: Continuous Action Q-LearningMoonkyung Ryu, Yinlam Chow, Ross Anderson, Christian Tjandraatmadja et al.ICLR 2020 · 50 citations
- DIMES: A Differentiable Meta Solver for Combinatorial Optimization ProblemsRuizhong Qiu, Zhiqing Sun, Yiming YangNeurIPS 2022 · 183 citations
- Brick-by-Brick: Combinatorial Construction with Deep Reinforcement LearningHyunsoo Chung, Jungtaek Kim, Boris Knyazev, Jinhwi Lee et al.NeurIPS 2021 · 29 citations
- GINO-Q: Learning an Asymptotically Optimal Index Policy for Restless Multi-armed BanditsGongpu Chen, Soung Chang Liew, Deniz GündüzAAAI 2026 · 1 citation
- Learning Multi-Timescale Abstractions for Hierarchical Combinatorial PlanningVivienne Huiling Wang, Tinghuai Wang, Joni PajarinenICML 2026
