Intra-agent speech permits zero-shot task acquisition
Chen Yan, Federico Carnevale, Petko Georgiev, Adam Santoro, Aurelia Guy, Alistair Muldal, Chia-Chun Hung, Josh Abramson, Timothy P. Lillicrap, Gregory Wayne
Abstract
Human language learners are exposed to a trickle of informative, context-sensitive language, but a flood of raw sensory data. Through both social language use and internal processes of rehearsal and practice, language learners are able to build high-level, semantic representations that explain their perceptions. Here, we take inspiration from such processes of "inner speech" in humans (Vygotsky, 1934) to better understand the role of intra-agent speech in embodied behaviour. First, we formally pose intra-agent speech as a semi-supervised problem and develop two algorithms that enable visually grounded captioning with little labeled language data. We then experimentally compute scaling curves over different amounts of labeled data and compare the data efficiency against a supervised learning baseline. Finally, we incorporate intra-agent speech into an embodied, mobile manipulator agent operating in a 3D virtual world, and show that with as few as 150 additional image captions, intra-agent speech endows the agent with the ability to manipulate and answer questions about a new object without any related task-directed experience (zero-shot). Taken together, our experiments suggest that modelling intra-agent speech is effective in enabling embodied agents to learn new tasks efficiently and without direct interaction experience. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0a7dcd7-b71e-4fab-add0-403116881890Cited by top-tier papers5
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus et al.ICML 2023 · 36 citations
- Using natural language and program abstractions to instill human inductive biases in machinesSreejan Kumar, Carlos G. Correa, Ishita Dasgupta, Raja Marjieh et al.NeurIPS 2022 · 34 citations
- Semantic HELM: A Human-Readable Memory for Reinforcement LearningFabian Paischer, Thomas Adler, Markus Hofmarcher, Sepp HochreiterNeurIPS 2023 · 21 citations
- Simple Embodied Language Learning as a Byproduct of Meta-Reinforcement LearningEvan Zheran Liu, Sahaana Suri, Tong Mu, Allan Zhou et al.ICML 2023 · 4 citations
- Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordinationRakshit S. Trivedi, Kartik Sharma, David C. ParkesNeurIPS 2025 · 3 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Semantic Exploration from Language Abstractions and Pretrained RepresentationsAllison C. Tam, Neil C. Rabinowitz, Andrew K. Lampinen, Nicholas A. Roy et al.NeurIPS 2022 · 85 citations
- LVLMs are Bad at Overhearing Human Referential CommunicationZhengxiang Wang, Weiling Li, Panagiotis Kaliosis, Owen Rambow et al.EMNLP 2025
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- Learning to Caption Images Through a Lifetime by Asking QuestionsTingke Shen, Amlan Kar, Sanja FidlerICCV 2019 · 33 citations
- Embodied Image Captioning: Self-Supervised Learning Agents for Spatially Coherent Image DescriptionsTommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio et al.ICCV 2025
