Learning Compositional Tasks from Language Instructions
Lajanugen Logeswaran, Wilka Carvalho, Honglak Lee
摘要
Systematic compositionality -the ability to combine learned knowledge and skills to solve novel tasks -is a key aspect of generalization in humans that allows us to understand and perform tasks described by novel language utterances. While progress has been made in supervised learning settings, no work has yet studied compositional generalization of a reinforcement learning agent following natural language instructions in an embodied environment. We develop a set of tasks in a photo-realistic simulated kitchen environment that allow us to study the degree to which a behavioral policy captures the systematicity in language by studying its zero-shot generalization performance on held out natural language instructions. We show that our agent which leverages a novel additive action-value decomposition in tandem with attention-based subgoal prediction is able to exploit composition in text instructions to generalize to unseen tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Exploring the Benefits of Training Expert Language Models over Instruction TuningJoel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim 等ICML 2023 · 被引用 97 次
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee 等ICLR 2023 · 被引用 2 次
- Watch Less, Do More: Implicit Skill Discovery for Video-Conditioned PolicyJiangxing Wang, Zongqing LuICLR 2025
它引用的顶会 Paper3
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
- Environmental drivers of systematicity and generalization in a situated agentFelix Hill, Andrew K. Lampinen, Rosalia Schneider, Stephen Clark 等ICLR 2020 · 被引用 109 次
- Good-Enough Compositional Data AugmentationJacob AndreasACL 2020 · 被引用 15 次
相关 Paper
- RLZero: Direct Policy Inference from Language Without In-Domain SupervisionHarshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli 等NeurIPS 2025 · 被引用 8 次
- CtD: Composition through Decomposition in Emergent CommunicationBoaz Carmeli, Ron Meir, Yonatan BelinkovICLR 2025
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement LearningValerie Chen, Abhinav Gupta, Kenneth MarinoICLR 2021 · 被引用 6 次
- Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement LearningHarry Zhao, Safa Alver, Harm van Seijen, Romain Laroche 等ICLR 2024 · 被引用 5 次
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 被引用 51 次
