Conceptual Reinforcement Learning for Language-Conditioned Tasks
Shaohui Peng, Xing Hu, Rui Zhang, Jiaming Guo, Qi Yi, Ruizhi Chen, Zidong Du, Ling Li, Qi Guo, Yunji Chen
摘要
Despite the broad application of deep reinforcement learning (RL), transferring and adapting the policy to unseen but similar environments is still a significant challenge. Recently, the language-conditioned policy is proposed to facilitate policy transfer through learning the joint representation of observation and text that catches the compact and invariant information across various environments. Existing studies of language-conditioned RL methods often learn the joint representation as a simple latent layer for the given instances (episode-specific observation and text), which inevitably includes noisy or irrelevant information and cause spurious correlations that are dependent on instances, thus hurting generalization performance and training efficiency. To address the above issue, we propose a conceptual reinforcement learning (CRL) framework to learn the concept-like joint representation for language-conditioned policy. The key insight is that concepts are compact and invariant representations in human cognition through extracting similarities from numerous instances in real-world. In CRL, we propose a multi-level attention encoder and two mutual information constraints for learning compact and invariant concepts. Verified in two challenging environments, RTFM and Messenger, CRL significantly improves the training efficiency (up to 70%) and generalization ability (up to 30%) to the new environment dynamics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Context Shift Reduction for Offline Meta-Reinforcement LearningYunkai Gao, Rui Zhang, Jiaming Guo, Fan Wu 等NeurIPS 2023 · 被引用 30 次
- Contrastive Modules with Temporal Attention for Multi-Task Reinforcement LearningSiming Lan, Rui Zhang, Qi Yi, Jiaming Guo 等NeurIPS 2023 · 被引用 18 次
- Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human AlignmentChen Zhang, Qiang He, Yuan Zhou, Elvis S. Liu 等ICML 2024 · 被引用 7 次
- Online Prototype Alignment for Few-shot Policy TransferQi Yi, Rui Zhang, Shaohui Peng, Jiaming Guo 等ICML 2023 · 被引用 5 次
- Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill LearningShaohui Peng, Xing Hu, Qi Yi, Rui Zhang 等AAAI 2024 · 被引用 4 次
它引用的顶会 Paper15
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 等ICML 2020 · 被引用 512 次
- Interactive Fiction Games: A Colossal AdventureMatthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi YuanAAAI 2020 · 被引用 242 次
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos 等ICML 2020 · 被引用 153 次
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 被引用 116 次
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal 等ICLR 2021 · 被引用 77 次
相关 Paper
- RTFM: Generalising to New Environment Dynamics via ReadingVictor Zhong, Tim Rocktäschel, Edward GrefenstetteICLR 2020 · 被引用 44 次
- Improving Policy Learning via Language Dynamics DistillationVictor Zhong, Jesse Mu, Luke Zettlemoyer, Edward Grefenstette 等NeurIPS 2022 · 被引用 16 次
- Natural Language Instruction-following with Task-related Language Development and TranslationJing-Cheng Pang, Xinyu Yang, Si-Hang Yang, Xiong-Hui Chen 等NeurIPS 2023 · 被引用 18 次
- Grounding Language to Entities and Dynamics for Generalization in Reinforcement LearningAustin W. Hanjie, Victor Zhong, Karthik NarasimhanICML 2021 · 被引用 60 次
- Zero-Shot Reward Specification via Grounded Natural LanguageParsa Mahmoudieh, Deepak Pathak, Trevor DarrellICML 2022 · 被引用 69 次
