Continuous-Time Value Iteration for Multi-Agent Reinforcement Learning
Xuefeng Wang, Lei Zhang, Henglin Pu, Ahmed Hussain Qureshi, Husheng Li
Abstract
Existing reinforcement learning (RL) methods struggle with complex dynamical systems that demand interactions at high frequencies or irregular time intervals. Continuous-time RL (CTRL) has emerged as a promising alternative by replacing discrete-time Bellman recursion with differentiable value functions defined as viscosity solutions of the Hamilton–Jacobi–Bellman (HJB) equation. While CTRL has shown promise, its applications have been largely limited to the single-agent domain. This limitation stems from two key challenges: (i) conventional methods for solving HJB equations suffer from the curse of dimensionality (CoD), making them intractable in high-dimensional systems; and (ii) even with learning-based approaches to alleviate the CoD, accurately approximating centralized value functions in multi-agent settings remains difficult, which in turn destabilizes policy training. In this paper, we propose a CT-MARL framework that uses physics-informed neural networks (PINNs) to approximate HJB-based value functions at scale. To ensure the value is consistent with its differential structure, we align value learning with value-gradient learning by introducing a Value Gradient Iteration (VGI) module that iteratively refines value gradients along trajectories. This improves gradient accuracy, in turn yielding more precise value approximations and stronger policy learning. We evaluate our method using continuous‑time variants of standard benchmarks, including multi‑agent particle environment (MPE) and multi‑agent MuJoCo. Our results demonstrate that our approach consistently outperforms existing continuous‑time RL baselines and scales to complex cooperative multi-agent dynamics. Code is available at https://github.com/Wangxuefeng1024/Continuous-Time-Value-Iteration-for-Multi-Agent-Reinforcement-Learning.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7f67a7e-1fa2-4890-9f83-e7c7d07d7058Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Characterizing possible failure modes in physics-informed neural networksAditi S. Krishnapriyan, Amir Gholami, Shandian Zhe, Robert M. Kirby et al.NeurIPS 2021 · 1,421 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 72 citations
- Time Discretization-Invariant Safe Action Repetition for Policy Gradient MethodsSeohong Park, Jaekyeom Kim, Gunhee KimNeurIPS 2021 · 33 citations
- Physics-Informed Neural Network Policy Iteration: Algorithms, Convergence, and VerificationYiming Meng, Ruikun Zhou, Amartya Mukherjee, Maxwell Fitzsimmons et al.ICML 2024 · 23 citations
Related papers
- Physics-Informed Approach for Exploratory Hamilton-Jacobi-Bellman Equations via Policy IterationsYeongjong Kim, Namkyeong Cho, Minseok Kim, Yeoneung KimAAAI 2026 · 4 citations
- A Physics-Informed Machine Learning Framework for Safe and Optimal Control of Autonomous SystemsManan Tayal, Aditya Singh, Shishir Kolathaya, Somil BansalICML 2025
- Deep learning for continuous-time stochastic control with jumpsPatrick Cheridito, Jean-Loup Dupret, Donatien HainautNeurIPS 2025 · 8 citations
- ConFIG: Towards Conflict-free Training of Physics Informed Neural NetworksQiang Liu, Mengyu Chu, Nils ThuereyICLR 2025
- Physics-informed Neural Networks for Functional Differential Equations: Cylindrical Approximation and Its Convergence GuaranteesTaiki Miyagawa, Takeru YokotaNeurIPS 2024 · 8 citations
