Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
Anirudh Goyal, Shagun Sodhani, Jonathan Binas, Xue Bin Peng, Sergey Levine, Yoshua Bengio
Abstract
Reinforcement learning agents that operate in diverse and complex environments can benefit from the structured decomposition of their behavior. Often, this is addressed in the context of hierarchical reinforcement learning, where the aim is to decompose a policy into lower-level primitives or options, and a higher-level meta-policy that triggers the appropriate behaviors for a given situation. However, the meta-policy must still produce appropriate decisions in all states. In this work, we propose a policy design that decomposes into primitives, similarly to hierarchical reinforcement learning, but without a high-level meta-policy. Instead, each primitive can decide for themselves whether they wish to act in the current state. We use an information-theoretic mechanism for enabling this decentralized decision: each primitive chooses how much information it needs about the current state to make a decision and the primitive that requests the most information about the current state acts in the world. The primitives are regularized to use as little information as possible, which leads to natural competition and specialization. We experimentally demonstrate that this policy architecture improves over both flat and hierarchical policies in terms of generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee21d238-c17e-4fea-838a-17ccc0ffcfd0Cited by top-tier papers14
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
- Multi-Task Reinforcement Learning with Soft ModularizationRuihan Yang, Huazhe Xu, Yi Wu, Xiaolong WangNeurIPS 2020 · 247 citations
- Learning to Coordinate Manipulation Skills via Skill Behavior DiversificationYoungwoon Lee, Jingyun Yang, Joseph J. LimICLR 2020 · 98 citations
- Retrieval-Augmented Reinforcement LearningAnirudh Goyal, Abram L. Friesen, Andrea Banino, Theophane Weber et al.ICML 2022 · 69 citations
- The Neural Race Reduction: Dynamics of Abstraction in Gated NetworksAndrew M. Saxe, Shagun Sodhani, Sam Jay LewallenICML 2022 · 52 citations
Related papers
- Flexible Option LearningMartin Klissarov, Doina PrecupNeurIPS 2021 · 38 citations
- RD: Reward Decomposition with Representation DecompositionZichuan Lin, Derek Yang, Li Zhao, Tao Qin et al.NeurIPS 2020 · 12 citations
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 97 citations
- Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized TeamingSachin G. Konan, Esmaeil Seraj, Matthew C. GombolayICLR 2022 · 27 citations
- OPtions as REsponses: Grounding behavioural hierarchies in multi-agent reinforcement learningAlexander Vezhnevets, Yuhuai Wu, Maria K. Eckstein, Rémi Leblond et al.ICML 2020 · 44 citations
