Recomposing the Reinforcement Learning Building Blocks with Hypernetworks
Elad Sarafian, Shai Keynan, Sarit Kraus
摘要
The Reinforcement Learning (RL) building blocks, i.e. Q-functions and policy networks, usually take elements from the cartesian product of two domains as input. In particular, the input of the Q-function is both the state and the action, and in multi-task problems (Meta-RL) the policy can take a state and a context. Standard architectures tend to ignore these variables' underlying interpretations and simply concatenate their features into a single vector. In this work, we argue that this choice may lead to poor gradient estimation in actor-critic algorithms and high variance learning steps in Meta-RL algorithms. To consider the interaction between the input variables, we suggest using a Hypernetwork architecture where a primary network determines the weights of a conditional dynamic network. We show that this approach improves the gradient approximation and reduces the learning step variance, which both accelerates learning and improves the final performance. We demonstrate a consistent improvement across different locomotion tasks and different algorithms both in RL (TD3 and SAC) and in Meta-RL (MAML and PEARL).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Learning a Diffusion Model Policy from Rewards via Q-Score MatchingMichael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi MaICML 2024 · 被引用 90 次
- PaCo: Parameter-Compositional Multi-task Reinforcement LearningLingfeng Sun, Haichao Zhang, Wei Xu, Masayoshi TomizukaNeurIPS 2022 · 被引用 72 次
- Dynamics Generalisation in Reinforcement Learning via Adaptive Context-Aware PoliciesMichael Beukman, Devon Jarvis, Richard Klein, Steven James 等NeurIPS 2023 · 被引用 29 次
- Universal Morphology Control via Contextual ModulationZheng Xiong, Jacob Beck, Shimon WhitesonICML 2023 · 被引用 27 次
- Controllable Dynamic Multi-Task ArchitecturesDripta S. Raychaudhuri, Yumin Suh, Samuel Schulter, Xiang Yu 等CVPR 2022 · 被引用 24 次
它引用的顶会 Paper10
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 被引用 412 次
- Meta-Q-LearningRasool Fakoor, Pratik Chaudhari, Stefano Soatto, Alexander J. SmolaICLR 2020 · 被引用 162 次
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz 等ICLR 2020 · 被引用 152 次
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 107 次
相关 Paper
- MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics ComponentsMin Wang, Xin Li, Leiji Zhang, Mingzhong WangAAAI 2024 · 被引用 6 次
- Online Meta-Critic Learning for Off-Policy Actor-Critic MethodsWei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang 等NeurIPS 2020 · 被引用 54 次
- Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement LearningXinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng LvKDD 2026
- Hypernetworks for Zero-Shot Transfer in Reinforcement LearningSahand Rezaei-Shoshtari, Charlotte Morissette, François Robert Hogan, Gregory Dudek 等AAAI 2023 · 被引用 23 次
- Mixture-of-World Models: Scaling Multi-Task Reinforcement Learning with Modular Latent DynamicsBoxuan Zhang, Weipu Zhang, Zhaohan Feng, Wei Xiao 等ICLR 2026 · 被引用 1 次
