A Benchmark for Systematic Generalization in Grounded Language Understanding
Laura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt, Brenden M. Lake
Abstract
Human language users easily interpret expressions that describe unfamiliar situations composed from familiar parts ("greet the pink brontosaurus by the ferris wheel"). Modern neural networks, by contrast, struggle to interpret compositions unseen in training. In this paper, we introduce a new benchmark, gSCAN, for evaluating compositional generalization in models of situated language understanding. We take inspiration from standard models of meaning composition in formal linguistics. Going beyond an earlier related benchmark that focused on syntactic aspects of generalization, gSCAN defines a language grounded in the states of a grid world. This allows us to build novel generalization tasks that probe the acquisition of linguistically motivated rules. For example, agents must understand how adjectives such as 'small' are interpreted relative to the current world state or how adverbs such as 'cautiously' combine with new verbs. We test a strong multi-modal baseline model and a state-of-the-art compositional method finding that, in most cases, they fail dramatically when generalization requires systematic compositional rules.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c35b293-f5bb-4f01-abcc-c94715a8fcdaCited by top-tier papers44
- Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic ReasoningMaxwell I. Nye, Michael Henry Tessler, Joshua B. Tenenbaum, Brenden M. LakeNeurIPS 2021 · 151 citations
- Compositional Generalization via Neural-Symbolic Stack MachinesXinyun Chen, Chen Liang, Adams Wei Yu, Dawn Song et al.NeurIPS 2020 · 112 citations
- Inducing Causal Structure for Interpretable Neural NetworksAtticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner et al.ICML 2022 · 104 citations
- Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language ModelsWenshan Wu, Shaoguang Mao, Yadong Zhang, Yan Xia et al.NeurIPS 2024 · 100 citations
- COLLIE: Systematic Construction of Constrained Text Generation TasksShunyu Yao, Howard Chen, Austin W. Hanjie, Runzhe Yang et al.ICLR 2024 · 65 citations
Builds on4
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- Learning Compositional Rules via Neural Program SynthesisMaxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. LakeNeurIPS 2020 · 120 citations
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 112 citations
- Good-Enough Compositional Data AugmentationJacob AndreasACL 2020 · 15 citations
Related papers
- When Can Transformers Ground and Compose: Insights from Compositional Generalization BenchmarksAnkur Sikarwar, Arkil Patel, Navin GoyalEMNLP 2022 · 4 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Meta-Learning to Compositionally GeneralizeHenry Conklin, Bailin Wang, Kenny Smith, Ivan TitovACL 2021
- Inducing Transformer's Compositional Generalization Ability via Auxiliary Sequence Prediction TasksYichen Jiang, Mohit BansalEMNLP 2021
- Compositional Generalization by Learning Analytical ExpressionsQian Liu, Shengnan An, Jian-Guang Lou, Bei Chen et al.NeurIPS 2020 · 79 citations
