Weakly Coupled Deep Q-Networks
Ibrahim El Shar, Daniel R. Jiang
Abstract
We propose weakly coupled deep Q-networks (WCDQN), a novel deep reinforcement learning algorithm that enhances performance in a class of structured problems called weakly coupled Markov decision processes (WCMDP). WCMDPs consist of multiple independent subproblems connected by an action space constraint, which is a structural property that frequently emerges in practice. Despite this appealing structure, WCMDPs quickly become intractable as the number of subproblems grows. WCDQN employs a single network to train multiple DQN "subagents," one for each subproblem, and then combine their solutions to establish an upper bound on the optimal action value. This guides the main DQN agent towards optimality. We show that the tabular version, weakly coupled Q-learning (WCQL), converges almost surely to the optimal action value. Numerical experiments show faster convergence compared to DQN and related techniques in settings with as many as 10 subproblems, 3 10 total actions, and a continuous state space. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f4f9f94-d16c-40eb-a41c-e1a329461097Cited by top-tier papers2
- Projection-based Lyapunov method for fully heterogeneous weakly-coupled MDPsXiangcheng Zhang, Yige Hong, Weina WangNeurIPS 2025 · 2 citations
- Multi-agent Markov EntanglementShuze Chen, Tianyi PengNeurIPS 2025
Builds on6
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 91 citations
- NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RLKhaled Nakhleh, Santosh Ganji, Ping-Chun Hsieh, I-Hong Hou et al.NeurIPS 2021 · 52 citations
- Learning Infinite-Horizon Average-Reward Restless Multi-Action Bandits via Index AwarenessGuojun Xiong, Shufan Wang, Jian LiNeurIPS 2022 · 21 citations
- Q-Learning Lagrange Policies for Multi-Action Restless BanditsJackson A. Killian, Arpita Biswas, Sanket Shah, Milind TambeKDD 2021 · 12 citations
- RL-MPCA: A Reinforcement Learning Based Multi-Phase Computation Allocation Approach for Recommender SystemsJiahong Zhou, Shunhui Mao, Guoliang Yang, Bo Tang et al.WWW 2023 · 10 citations
Related papers
- CAQL: Continuous Action Q-LearningMoonkyung Ryu, Yinlam Chow, Ross Anderson, Christian Tjandraatmadja et al.ICLR 2020 · 50 citations
- Convergent and Efficient Deep Q Learning AlgorithmZhikang T. Wang, Masahito UedaICLR 2022 · 3 citations
- Semi-infinitely Constrained Markov Decision ProcessesLiangyu Zhang, Yang Peng, Wenhao Yang, Zhihua ZhangNeurIPS 2022 · 5 citations
- Deep Coordination GraphsWendelin Boehmer, Vitaly Kurin, Shimon WhitesonICML 2020 · 209 citations
- Inferring DQN structure for high-dimensional continuous controlAndrey Sakryukin, Chedy Raïssi, Mohan S. KankanhalliICML 2020 · 8 citations
