Flexible Attention-Based Multi-Policy Fusion for Efficient Deep Reinforcement Learning
Zih-Yun Chiu, Yi-Lin Tuan, William Yang Wang, Michael C. Yip
Abstract
Reinforcement learning (RL) agents have long sought to approach the efficiency of human learning. Humans are great observers who can learn by aggregating external knowledge from various sources, including observations from others' policies of attempting a task. Prior studies in RL have incorporated external knowledge policies to help agents improve sample efficiency. However, it remains non-trivial to perform arbitrary combinations and replacements of those policies, an essential feature for generalization and transferability. In this work, we present Knowledge-Grounded RL (KGRL), an RL paradigm fusing multiple knowledge policies and aiming for human-like efficiency and flexibility. We propose a new actor architecture for KGRL, Knowledge-Inclusive Attention Network (KIAN), which allows free knowledge rearrangement due to embedding-based attentive action prediction. KIAN also addresses entropy imbalance, a problem arising in maximum entropy KGRL that hinders an agent from efficiently exploring the environment, through a new design of policy distributions. The experimental results demonstrate that KIAN outperforms alternative methods incorporating external knowledge policies and achieves efficient and flexible learning. Our implementation is available at https://github.com/Pascalson/KGRL.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98e7e1b8-4b51-41b5-84d8-b2aca5e7f360Builds on5
- Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate ProgressRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2022 · 95 citations
- Reinforcement Learning Under Moral UncertaintyAdrien Ecoffet, Joel LehmanICML 2021 · 40 citations
- Unsupervised Skill Discovery with Bottleneck Option LearningJaekyeom Kim, Seohong Park, Gunhee KimICML 2021 · 39 citations
- Composing Task-Agnostic Policies with Deep Reinforcement LearningAhmed Hussain Qureshi, Jacob J. Johnson, Yuzhe Qin, Taylor Henderson et al.ICLR 2020 · 35 citations
- Toward Robust Long Range Policy TransferWei-Cheng Tseng, Jin-Siang Lin, Yao-Min Feng, Min SunAAAI 2021 · 8 citations
Related papers
- PAE: Reinforcement Learning from External Knowledge for Efficient ExplorationZhe Wu, Haofei Lu, Junliang Xing, You Wu et al.ICLR 2024 · 1 citation
- Learning Multi-Agent Communication through Structured Attentive ReasoningMurtaza Rangwala, Ryan WilliamsNeurIPS 2020 · 42 citations
- Case-based reasoning for better generalization in textual reinforcement learningMattia Atzeni, Shehzaad Zuzar Dhuliawala, Keerthiram Murugesan, Mrinmaya SachanICLR 2022 · 16 citations
- Multi-Agent Actor-Critic with Hierarchical Graph Attention NetworkHeechang Ryu, Hayong Shin, Jinkyoo ParkAAAI 2020 · 143 citations
- Knowledge-Driven Virtual Network Embedding with Dynamic World ModelYangzi Song, Baoquan Ren, Yulong Shen, Qijie Qian et al.INFOCOM 2026
