Deep Radial-Basis Value Functions for Continuous Control
Kavosh Asadi, Neev Parikh, Ronald E. Parr, George Dimitri Konidaris, Michael L. Littman
摘要
A core operation in reinforcement learning (RL) is finding an action that is optimal with respect to a learned value function. This operation is often challenging when the learned value function takes continuous actions as input. We introduce deep radial-basis value functions (RBVFs): value functions learned using a deep network with a radial-basis function (RBF) output layer. We show that the maximum action-value with respect to a deep RBVF can be approximated easily and accurately. Moreover, deep RBVFs can represent any true value function owing to their support for universal function approximation. We extend the standard DQN algorithm to continuous control by endowing the agent with a deep RBVF. We show that the resultant agent, called RBF-DQN, significantly outperforms value-function-only baselines, and is competitive with state-of-the-art actor-critic algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Learning Markov State Abstractions for Deep Reinforcement LearningCameron Allen, Neev Parikh, Omer Gottesman, George KonidarisNeurIPS 2021 · 被引用 66 次
- Resetting the Optimizer in Deep RL: An Empirical StudyKavosh Asadi, Rasool Fakoor, Shoham SabachNeurIPS 2023 · 被引用 38 次
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin 等ICML 2022 · 被引用 32 次
- Optimistic Initialization for Exploration in Continuous ControlSam Lobel, Omer Gottesman, Cameron Allen, Akhil Bagaria 等AAAI 2022 · 被引用 14 次
- Q-functionals for Value-Based Continuous ControlSamuel Lobel, Sreehari Rammohan, Bowen He, Shangqun Yu 等AAAI 2023 · 被引用 10 次
它引用的顶会 Paper1
相关 Paper
- Learning State Representations from Random Deep Action-conditional PredictionsZeyu Zheng, Vivek Veeriah, Risto Vuorio, Richard L. Lewis 等NeurIPS 2021 · 被引用 6 次
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 被引用 25 次
- Actor-Free Continuous Control via Structurally Maximizable Q-FunctionsYigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem BiyikNeurIPS 2025
- Parameter-Based Value FunctionsFrancesco Faccio, Louis Kirsch, Jürgen SchmidhuberICLR 2021 · 被引用 29 次
- Continuous Deep Q-Learning in Optimal Control Problems: Normalized Advantage Functions AnalysisAnton Plaksin, Stepan MartyanovNeurIPS 2022 · 被引用 5 次
