Learning the Target Network in Function Space
Kavosh Asadi, Yao Liu, Shoham Sabach, Ming Yin, Rasool Fakoor
Abstract
We focus on the task of learning the value function in the reinforcement learning (RL) setting. This task is often solved by updating a pair of online and target networks while ensuring that the parameters of these two networks are equivalent. We propose Lookahead-Replicate (LR), a new valuefunction approximation algorithm that is agnostic to this parameter-space equivalence. Instead, the LR algorithm is designed to maintain an equivalence between the two networks in the function space. This value-based equivalence is obtained by employing a new target-network update. We show that LR leads to a convergent behavior in learning the value function. We also present empirical results demonstrating that LR-based targetnetwork updates significantly improve deep RL on the Atari benchmark. * Equal contribution. The order of the first authors was fully decided by dice roll.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 434b52a3-3448-4999-97d7-ba65057ff4a2Cited by top-tier papers1
Ask how each one uses itBuilds on8
- A Finite-Time Analysis of Two Time-Scale Actor-Critic MethodsYue Wu, Weitong Zhang, Pan Xu, Quanquan GuNeurIPS 2020 · 189 citations
- Towards Instance-Optimal Offline Reinforcement Learning with PessimismMing Yin, Yu-Xiang WangNeurIPS 2021 · 93 citations
- Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with PessimismMing Yin, Yaqi Duan, Mengdi Wang, Yu-Xiang WangICLR 2022 · 74 citations
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
- Resetting the Optimizer in Deep RL: An Empirical StudyKavosh Asadi, Rasool Fakoor, Shoham SabachNeurIPS 2023 · 38 citations
Related papers
- Faster Deep Reinforcement Learning with Slower Online NetworkKavosh Asadi, Rasool Fakoor, Omer Gottesman, Taesup Kim et al.NeurIPS 2022 · 7 citations
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement LearningAhmed Hendawy, Henrik Metternich, Théo Vincent, Mahdi Kallel et al.ICLR 2026 · 4 citations
- A new convergent variant of Q-learning with linear function approximationDiogo S. Carvalho, Francisco S. Melo, Pedro SantosNeurIPS 2020 · 39 citations
- Frustratingly Easy Regularization on Representation Can Boost Deep Reinforcement LearningQiang He, Huangyuan Su, Jieyu Zhang, Xinwen HouCVPR 2023
- Bridging the performance-gap between target-free and target-based reinforcement learningThéo Vincent, Yogesh Tripathi, Tim Lukas Faust, Abdullah Akgül et al.ICLR 2026 · 6 citations
