Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks
Fabio Pardo, Vitaly Levdik, Petar Kormushev
Abstract
Being able to reach any desired location in the environment can be a valuable asset for an agent. Learning a policy to navigate between all pairs of states individually is often not feasible. An all-goals updating algorithm uses each transition to learn Q-values towards all goals simultaneously and off-policy. However the expensive numerous updates in parallel limited the approach to small tabular cases so far. To tackle this problem we propose to use convolutional network architectures to generate Q-values and updates for a large number of goals at once. We demonstrate the accuracy and generalization qualities of the proposed method on randomly generated mazes and Sokoban puzzles. In the case of on-screen goal coordinates the resulting mapping from frames to distance-maps directly informs the agent about which places are reachable and in how many steps. As an example of application we show that replacing the random actions in ε-greedy exploration by several actions towards feasible goals generates better exploratory trajectories on Montezuma's Revenge and Super Mario All-Stars games.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Goal-Conditioned Agents that Learn Everything All at OnceMichael Matthews, Matthew Jackson, Michael Beukman, Thomas Foster et al.ICML 2026
- Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNNMohammad Taufeeque, Aaron David Tucker, Adam Gleave, Adrià Garriga-AlonsoICLR 2026 · 5 citations
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 11 citations
- PlanGAN: Model-based Planning With Sparse Rewards and Multiple GoalsHenry Charlesworth, Giovanni MontanaNeurIPS 2020 · 34 citations
- Goal-Conditioned Q-learning as Knowledge DistillationAlexander Levine, Soheil FeiziAAAI 2023 · 4 citations
