Continuous Soft Actor-Critic: An Off-Policy Learning Method Robust to Time Discretization
Huimin Han, Shaolin Ji
摘要
Many Deep Reinforcement Learning (DRL) algorithms are sensitive to time discretization, which reduces their performance in real-world scenarios. We propose Continuous Soft Actor-Critic, an off-policy actor-critic DRL algorithm in continuous time and space. It is robust to environment time discretization. We also extend the framework to multi-agent scenarios. This Multi-Agent Reinforcement Learning (MARL) algorithm is suitable for both competitive and cooperative settings. Policy evaluation employs stochastic control theory, with loss functions derived from martingale orthogonality conditions. We establish scaling principles for hyper-parameters of the algorithm as the environment time discretization δt changes ( δt → 0 ). We provide theoretical proofs for the relevant theorems. To validate the algorithm’s effectiveness, we conduct comparative experiments between the proposed algorithm and other mainstream methods across multiple tasks in Virtual Multi-Agent System (VMAS). Experimental results demonstrate that the proposed algorithm achieves robust performance across various environments with different time discretization parameter settings, outperforming other methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Towards a Standardised Performance Evaluation Protocol for Cooperative MARLRihab Gorsane, Omayma Mahjoub, Ruan de Kock, Roland Dubb 等NeurIPS 2022 · 被引用 79 次
- TorchRL: A data-driven decision-making library for PyTorchAlbert Bou, Matteo Bettini, Sebastian Dittert, Vikash Kumar 等ICLR 2024 · 被引用 77 次
相关 Paper
- From Ticks to Flows: Dynamics of Neural Reinforcement Learning in Continuous EnvironmentsSaket Tiwari, Tejas Kotwal, George Dimitri KonidarisICLR 2026
- Solving Continuous Control via Q-learningTim Seyde, Peter Werner, Wilko Schwarting, Igor Gilitschenski 等ICLR 2023 · 被引用 3 次
- Multi-Agent Reinforcement Learning in Stochastic Networked SystemsYiheng Lin, Guannan Qu, Longbo Huang, Adam WiermanNeurIPS 2021 · 被引用 55 次
- Reinforcement Learning with Random DelaysYann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher J. Pal 等ICLR 2021 · 被引用 3 次
- Divergence-Regularized Multi-Agent Actor-CriticKefan Su, Zongqing LuICML 2022 · 被引用 31 次
