Continuous Soft Actor-Critic: An Off-Policy Learning Method Robust to Time Discretization
Huimin Han, Shaolin Ji
Abstract
Many Deep Reinforcement Learning (DRL) algorithms are sensitive to time discretization, which reduces their performance in real-world scenarios. We propose Continuous Soft Actor-Critic, an off-policy actor-critic DRL algorithm in continuous time and space. It is robust to environment time discretization. We also extend the framework to multi-agent scenarios. This Multi-Agent Reinforcement Learning (MARL) algorithm is suitable for both competitive and cooperative settings. Policy evaluation employs stochastic control theory, with loss functions derived from martingale orthogonality conditions. We establish scaling principles for hyper-parameters of the algorithm as the environment time discretization δt changes ( δt → 0 ). We provide theoretical proofs for the relevant theorems. To validate the algorithm’s effectiveness, we conduct comparative experiments between the proposed algorithm and other mainstream methods across multiple tasks in Virtual Multi-Agent System (VMAS). Experimental results demonstrate that the proposed algorithm achieves robust performance across various environments with different time discretization parameter settings, outperforming other methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f981b2e3-a99e-4822-bb46-ed5e66e15f75Builds on3
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb et al.NeurIPS 2022 · 79 citations
- TorchRL: A data-driven decision-making library for PyTorchAlbert Bou, Matteo Bettini, Sebastian Dittert, Vikash Kumar et al.ICLR 2024 · 77 citations
Related papers
- From Ticks to Flows: Dynamics of Neural Reinforcement Learning in Continuous EnvironmentsSaket Tiwari, Tejas Kotwal, George Dimitri KonidarisICLR 2026
- Solving Continuous Control via Q-learningTim Seyde, Peter Werner, Wilko Schwarting, Igor Gilitschenski et al.ICLR 2023 · 3 citations
- Multi-Agent Reinforcement Learning in Stochastic Networked SystemsYiheng Lin, Guannan Qu, Longbo Huang, Adam WiermanNeurIPS 2021 · 55 citations
- Reinforcement Learning with Random DelaysYann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal et al.ICLR 2021 · 3 citations
- Divergence-Regularized Multi-Agent Actor-CriticKefan Su, Zongqing LuICML 2022 · 31 citations
