Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
Zequan Wu, Mengye Ren
Abstract
The Forward-Forward (FF) Algorithm is a recently proposed learning procedure for neural networks that employs two forward passes instead of the traditional forward and backward passes used in backpropagation. However, FF remains largely confined to supervised settings, leaving a gap at domains where learning signals can be yielded more naturally such as RL. In this work, inspired by FF's goodness function using layer activity statistics, we introduce Action-conditioned Root mean squared Q-Functions (ARQ), a novel value estimation method that applies a goodness function and action conditioning for local RL using temporal difference learning. Despite its simplicity and biological grounding, our approach achieves superior performance compared to state-of-the-art local backprop-free RL methods in the MinAtar and the DeepMind Control Suite benchmarks, while also outperforming algorithms trained with backpropagation on most tasks. Code can be found at https://github.com/agentic-learning-ai-lab/arq .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f2eb6fa-0e0a-4ffb-92bd-f1b10bd6e5b2Builds on14
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
Related papers
- Temporal-Difference Learning Using Distributed Error SignalsJonas Guan, Shon Eduard Verch, Claas Voelcker, Ethan C. Jackson et al.NeurIPS 2024 · 5 citations
- Stochastic Forward-Forward Learning through Representational Dimensionality CompressionZhichao Zhu, Yang Qi, Hengyuan Ma, Wenlian Lu et al.NeurIPS 2025 · 4 citations
- Learning the Target Network in Function SpaceKavosh Asadi, Yao Liu, Shoham Sabach, Ming Yin et al.ICML 2024 · 3 citations
- Accelerating Q-learning through Efficient Value-sharing across ActionsPrabhat Nagarajan, Brett Daley, Martha White, Marlos C. MachadoICML 2026
- DeeperForward: Enhanced Forward-Forward Training for Deeper and Better PerformanceLiang Sun, Yang Zhang, Weizhao He, Jiajun Wen et al.ICLR 2025
