Distilling Internet-Scale Vision-Language Models into Embodied Agents
Theodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus, Ishita Dasgupta
摘要
Instruction-following agents must ground language into their observation and action spaces. Yet learning to ground language is challenging, typically requiring domain-specific engineering or large quantities of human interaction data. To address this challenge, we propose using pretrained vision-language models (VLMs) to supervise embodied agents. We combine ideas from model distillation and hindsight experience replay (HER), using a VLM to retroactively generate language describing the agent's behavior. Simple prompting allows us to control the supervision signal, teaching an agent to interact with novel objects based on their names (e.g., planes) or their features (e.g., colors) in a 3D rendered environment. Fewshot prompting lets us teach abstract category membership, including pre-existing categories (food vs toys) and ad-hoc ones (arbitrary preferences over objects). Our work outlines a new and effective way to use internet-scale VLMs, repurposing the generic language grounding acquired by such models to teach task-relevant groundings to embodied agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez 等ICLR 2024 · 被引用 154 次
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba 等NeurIPS 2023 · 被引用 123 次
- BAGEL: Bootstrapping Agents by Guiding Exploration with LanguageShikhar Murty, Christopher D. Manning, Peter Shaw, Mandar Joshi 等ICML 2024 · 被引用 33 次
- FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement LearningYuwei Fu, Haichao Zhang, Di Wu, Wei Xu 等ICML 2024 · 被引用 31 次
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao 等CVPR 2024 · 被引用 23 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
相关 Paper
- Bridging Environments and Language with Rendering Functions and Vision-Language ModelsThéo Cachet, Christopher R. Dance, Olivier SigaudICML 2024 · 被引用 1 次
- MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDYZiyuan Qin, Huahui Yi, Qicheng Lao, Kang LiICLR 2023 · 被引用 25 次
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationCheng-Fu Yang, Da Yin, Wenbo Hu, Heng Ji 等ICCV 2025 · 被引用 1 次
- VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement LearningHaoran Xu, Peixi Peng, Guang Tan, Yiqian Chang 等CVPR 2025
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger 等NeurIPS 2023 · 被引用 63 次
