PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World
Rowan Zellers, Ari Holtzman, Matthew E. Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, Yejin Choi
摘要
We propose PIGLeT: a model that learns physical commonsense knowledge through interaction, and then uses this knowledge to ground language. We factorize PIGLeT into a physical dynamics model, and a separate language model. Our dynamics model learns not just what objects are but also what they do: glass cups break when thrown, plastic ones don't. We then use it as the interface to our language model, giving us a unified model of linguistic form and grounded meaning. PIGLeT can read a sentence, simulate neurally what might happen next, and then communicate that result through a literal symbolic representation, or natural language. Experimental results show that our model effectively learns world dynamics, along with how to communicate them. It is able to correctly forecast "what happens next" given an English sentence over 80% of the time, outperforming a 100x larger, text-to-text approach by over 10%. Likewise, its natural language summaries of physical interactions are also judged by humans as more accurate than LM alternatives. We present comprehensive analysis showing room for future work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Language Models Meet World Models: Embodied Experiences Enhance Language ModelsJiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu 等NeurIPS 2023 · 被引用 180 次
- Skill Induction and Planning with Latent LanguagePratyusha Sharma, Antonio Torralba, Jacob AndreasACL 2022 · 被引用 127 次
- Large Language Models Are Neurosymbolic ReasonersMeng Fang, Shilong Deng, Yudi Zhang, Zijing Shi 等AAAI 2024 · 被引用 53 次
- A Data Source for Reasoning Embodied AgentsJack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve, Yuxuan Sun 等AAAI 2023 · 被引用 10 次
- Learning Planning Abstractions from LanguageWeiyu Liu, Geng Chen, Joy Hsu, Jiayuan Mao 等ICLR 2024 · 被引用 6 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 被引用 914 次
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt 等NeurIPS 2020 · 被引用 169 次
- Experience Grounds LanguageYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas 等EMNLP 2020 · 被引用 74 次
相关 Paper
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear 等ICML 2020 · 被引用 88 次
- Learning to Model the World With LanguageJessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner 等ICML 2024 · 被引用 76 次
- Can Vision Language Models Learn Intuitive Physics from Interaction?Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan, Eric SchulzICML 2026
- Mind's Eye: Grounded Language Model Reasoning through SimulationRuibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu 等ICLR 2023 · 被引用 22 次
- SILG: The Multi-domain Symbolic Interactive Language Grounding BenchmarkVictor Zhong, Austin W. Hanjie, Sida I. Wang, Karthik Narasimhan 等NeurIPS 2021 · 被引用 23 次
