Continuous Deep Q-Learning in Optimal Control Problems: Normalized Advantage Functions Analysis
Anton Plaksin, Stepan Martyanov
Abstract
One of the most effective continuous deep reinforcement learning algorithms is normalized advantage functions (NAF). The main idea of NAF consists in the approximation of the Q-function by functions quadratic with respect to the action variable. This idea allows to apply the algorithm to continuous reinforcement learning problems, but on the other hand, it brings up the question of classes of problems in which this approximation is acceptable. The presented paper describes one such class. We consider reinforcement learning problems obtained by the time-discretization of certain optimal control problems. Based on the idea of NAF, we present a new family of quadratic functions and prove its suitable approximation properties. Taking these properties into account, we provide several ways to improve NAF. The experimental results confirm the efficiency of our improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cabfaba-290e-4f9a-bdf2-1b2b35f05a2bBuilds on2
Related papers
- Deep Radial-Basis Value Functions for Continuous ControlKavosh Asadi, Neev Parikh, Ronald E. Parr, George Dimitri Konidaris et al.AAAI 2021 · 25 citations
- Q-functionals for Value-Based Continuous ControlSamuel Lobel, Sreehari Rammohan, Bowen He, Shangqun Yu et al.AAAI 2023 · 10 citations
- Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time ControlChengxiu Hua, Jiawen Gu, Yushun TangNeurIPS 2025 · 5 citations
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin et al.ICML 2022 · 32 citations
- Zap Q-Learning With Nonlinear Function ApproximationShuhang Chen, Adithya M. Devraj, Fan Lu, Ana Busic et al.NeurIPS 2020 · 26 citations
