Continuous Deep Q-Learning in Optimal Control Problems: Normalized Advantage Functions Analysis
Anton Plaksin, Stepan Martyanov
摘要
One of the most effective continuous deep reinforcement learning algorithms is normalized advantage functions (NAF). The main idea of NAF consists in the approximation of the Q-function by functions quadratic with respect to the action variable. This idea allows to apply the algorithm to continuous reinforcement learning problems, but on the other hand, it brings up the question of classes of problems in which this approximation is acceptable. The presented paper describes one such class. We consider reinforcement learning problems obtained by the time-discretization of certain optimal control problems. Based on the idea of NAF, we present a new family of quadratic functions and prove its suitable approximation properties. Taking these properties into account, we provide several ways to improve NAF. The experimental results confirm the efficiency of our improvements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Deep Radial-Basis Value Functions for Continuous ControlKavosh Asadi, Neev Parikh, Ronald E. Parr, George Dimitri Konidaris 等AAAI 2021 · 被引用 25 次
- Q-functionals for Value-Based Continuous ControlSamuel Lobel, Sreehari Rammohan, Bowen He, Shangqun Yu 等AAAI 2023 · 被引用 10 次
- Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time ControlChengxiu Hua, Jiawen Gu, Yushun TangNeurIPS 2025 · 被引用 5 次
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin 等ICML 2022 · 被引用 32 次
- Zap Q-Learning With Nonlinear Function ApproximationShuhang Chen, Adithya M. Devraj, Fan Lu, Ana Busic 等NeurIPS 2020 · 被引用 26 次
