Deep Radial-Basis Value Functions for Continuous Control
Kavosh Asadi, Neev Parikh, Ronald E. Parr, George Dimitri Konidaris, Michael L. Littman
Abstract
A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value functions learned using a deep network with a radial-basis function (RBF) output layer. We show that the maximum action-value with respect to a deep RBVF can be approximated easily and accurately. Moreover, deep RBVFs can represent any true value function owing to their support for universal function approximation. We extend the standard DQN algorithm to continuous control by endowing the agent with a deep RBVF. We show that the resultant agent, called RBF-DQN, significantly outperforms value-function-only baselines, and is competitive with state-of-the-art actor-critic algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6af0f7c-8350-4db6-a993-4de4a292e1caCited by top-tier papers8
- Learning Markov State Abstractions for Deep Reinforcement LearningCameron Allen, Neev Parikh, Omer Gottesman, George KonidarisNeurIPS 2021 · 66 citations
- Resetting the Optimizer in Deep RL: An Empirical StudyKavosh Asadi, Rasool Fakoor, Shoham SabachNeurIPS 2023 · 38 citations
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin et al.ICML 2022 · 32 citations
- Optimistic Initialization for Exploration in Continuous ControlSam Lobel, Omer Gottesman, Cameron Allen, Akhil Bagaria et al.AAAI 2022 · 14 citations
- Q-functionals for Value-Based Continuous ControlSamuel Lobel, Sreehari Rammohan, Bowen He, Shangqun Yu et al.AAAI 2023 · 10 citations
Builds on1
Related papers
- Learning State Representations from Random Deep Action-conditional PredictionsZeyu Zheng, Vivek Veeriah, Risto Vuorio, Richard L. Lewis et al.NeurIPS 2021 · 6 citations
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 25 citations
- Actor-Free Continuous Control via Structurally Maximizable Q-FunctionsYigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem BiyikNeurIPS 2025
- Parameter-Based Value FunctionsFrancesco Faccio, Louis Kirsch, Jürgen SchmidhuberICLR 2021 · 29 citations
- Continuous Deep Q-Learning in Optimal Control Problems: Normalized Advantage Functions AnalysisAnton Plaksin, Stepan MartyanovNeurIPS 2022 · 5 citations
