PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World
Rowan Zellers, Ari Holtzman, Matthew E. Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, Yejin Choi
Abstract
We propose PIGLeT: a model that learns physical commonsense knowledge through interaction, and then uses this knowledge to ground language. We factorize PIGLeT into a physical dynamics model, and a separate language model. Our dynamics model learns not just what objects are but also what they do: glass cups break when thrown, plastic ones don't. We then use it as the interface to our language model, giving us a unified model of linguistic form and grounded meaning. PIGLeT can read a sentence, simulate neurally what might happen next, and then communicate that result through a literal symbolic representation, or natural language. Experimental results show that our model effectively learns world dynamics, along with how to communicate them. It is able to correctly forecast "what happens next" given an English sentence over 80% of the time, outperforming a 100x larger, text-to-text approach by over 10%. Likewise, its natural language summaries of physical interactions are also judged by humans as more accurate than LM alternatives. We present comprehensive analysis showing room for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfdf7ffa-3a86-4b9b-a7ed-86ed4f2c00aaCited by top-tier papers14
- Language Models Meet World Models: Embodied Experiences Enhance Language ModelsJiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu et al.NeurIPS 2023 · 180 citations
- Skill Induction and Planning with Latent LanguagePratyusha Sharma, Antonio Torralba, Jacob AndreasACL 2022 · 127 citations
- Large Language Models Are Neurosymbolic ReasonersMeng Fang, Shilong Deng, Yudi Zhang, Zijing Shi et al.AAAI 2024 · 53 citations
- A Data Source for Reasoning Embodied AgentsJack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve, Yuxuan Sun et al.AAAI 2023 · 10 citations
- Learning Planning Abstractions from LanguageWeiyu Liu, Geng Chen, Joy Hsu, Jiayuan Mao et al.ICLR 2024 · 6 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
- Experience Grounds LanguageYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas et al.EMNLP 2020 · 74 citations
Related papers
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear et al.ICML 2020 · 88 citations
- Learning to Model the World With LanguageJessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner et al.ICML 2024 · 76 citations
- Can Vision Language Models Learn Intuitive Physics from Interaction?Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan, Eric SchulzICML 2026
- Mind's Eye: Grounded Language Model Reasoning through SimulationRuibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu et al.ICLR 2023 · 22 citations
- SILG: The Multi-domain Symbolic Interactive Language Grounding BenchmarkVictor Zhong, Austin W. Hanjie, Sida I. Wang, Karthik Narasimhan et al.NeurIPS 2021 · 23 citations
