Conceptual Reinforcement Learning for Language-Conditioned Tasks
Shaohui Peng, Xing Hu, Rui Zhang, Jiaming Guo, Qi Yi, Ruizhi Chen, Zidong Du, Ling Li, Qi Guo, Yunji Chen
Abstract
Despite the broad application of deep reinforcement learning (RL), transferring and adapting the policy to unseen but similar environments is still a significant challenge. Recently, the language-conditioned policy is proposed to facilitate policy transfer through learning the joint representation of observation and text that catches the compact and invariant information across various environments. Existing studies of language-conditioned RL methods often learn the joint representation as a simple latent layer for the given instances (episode-specific observation and text), which inevitably includes noisy or irrelevant information and cause spurious correlations that are dependent on instances, thus hurting generalization performance and training efficiency. To address the above issue, we propose a conceptual reinforcement learning (CRL) framework to learn the concept-like joint representation for language-conditioned policy. The key insight is that concepts are compact and invariant representations in human cognition through extracting similarities from numerous instances in real-world. In CRL, we propose a multi-level attention encoder and two mutual information constraints for learning compact and invariant concepts. Verified in two challenging environments, RTFM and Messenger, CRL significantly improves the training efficiency (up to 70%) and generalization ability (up to 30%) to the new environment dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e00a5b9-3896-4089-a7a5-75032ce59713Cited by top-tier papers6
- Context Shift Reduction for Offline Meta-Reinforcement LearningYunkai Gao, Rui Zhang, Jiaming Guo, Fan Wu et al.NeurIPS 2023 · 30 citations
- Contrastive Modules with Temporal Attention for Multi-Task Reinforcement LearningSiming Lan, Rui Zhang, Qi Yi, Jiaming Guo et al.NeurIPS 2023 · 18 citations
- Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human AlignmentChen Zhang, Qiang He, Yuan Zhou, Elvis S. Liu et al.ICML 2024 · 7 citations
- Online Prototype Alignment for Few-shot Policy TransferQi Yi, Rui Zhang, Shaohui Peng, Jiaming Guo et al.ICML 2023 · 5 citations
- Hypothesis, Verification, and Induction: Grounding Large Language Models with Self-Driven Skill LearningShaohui Peng, Xing Hu, Qi Yi, Rui Zhang et al.AAAI 2024 · 4 citations
Builds on15
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Interactive Fiction Games: A Colossal AdventureMatthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi YuanAAAI 2020 · 242 citations
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos et al.ICML 2020 · 153 citations
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 116 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
Related papers
- RTFM: Generalising to New Environment Dynamics via ReadingVictor Zhong, Tim Rocktäschel, Edward GrefenstetteICLR 2020 · 44 citations
- Improving Policy Learning via Language Dynamics DistillationVictor Zhong, Jesse Mu, Luke Zettlemoyer, Edward Grefenstette et al.NeurIPS 2022 · 16 citations
- Natural Language Instruction-following with Task-related Language Development and TranslationJing-Cheng Pang, Xinyu Yang, Si-Hang Yang, Xiong-Hui Chen et al.NeurIPS 2023 · 18 citations
- Grounding Language to Entities and Dynamics for Generalization in Reinforcement LearningAustin W. Hanjie, Victor Zhong, Karthik NarasimhanICML 2021 · 60 citations
- Zero-Shot Reward Specification via Grounded Natural LanguageParsa Mahmoudieh, Deepak Pathak, Trevor DarrellICML 2022 · 69 citations
