Lune

NeurIPS2025Top-tier venue

Continuous Soft Actor-Critic: An Off-Policy Learning Method Robust to Time Discretization

Huimin Han, Shaolin Ji

2025Year

Abstract

Many Deep Reinforcement Learning (DRL) algorithms are sensitive to time discretization, which reduces their performance in real-world scenarios. We propose Continuous Soft Actor-Critic, an off-policy actor-critic DRL algorithm in continuous time and space. It is robust to environment time discretization. We also extend the framework to multi-agent scenarios. This Multi-Agent Reinforcement Learning (MARL) algorithm is suitable for both competitive and cooperative settings. Policy evaluation employs stochastic control theory, with loss functions derived from martingale orthogonality conditions. We establish scaling principles for hyper-parameters of the algorithm as the environment time discretization δt changes ( δt → 0 ). We provide theoretical proofs for the relevant theorems. To validate the algorithm’s effectiveness, we conduct comparative experiments between the proposed algorithm and other mainstream methods across multiple tasks in Virtual Multi-Agent System (VMAS). Experimental results demonstrate that the proposed algorithm achieves robust performance across various environments with different time discretization parameter settings, outperforming other methods.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f981b2e3-a99e-4822-bb46-ed5e66e15f75

Builds on3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines