COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
Najoung Kim, Tal Linzen
Abstract
Natural language is characterized by compositionality: the meaning of a complex expression is constructed from the meanings of its constituent parts. To facilitate the evaluation of the compositional abilities of language processing architectures, we introduce COGS, a semantic parsing dataset based on a fragment of English. The evaluation portion of COGS contains multiple systematic gaps that can only be addressed by compositional generalization; these include new combinations of familiar syntactic structures, or new combinations of familiar words and familiar structures. In experiments with Transformers and LSTMs, we found that in-distribution accuracy on the COGS test set was near-perfect (96-99%), but generalization accuracy was substantially lower (16-35%) and showed high sensitivity to random seed (±6-8%). These findings indicate that contemporary standard NLP models are limited in their compositional generalization capacity, and position COGS as a good way to measure progress.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ce726ae-74af-4271-94a9-d4cfc9eb5be6Cited by top-tier papers96
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao et al.ICLR 2022 · 237 citations
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 87 citations
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang et al.NeurIPS 2025 · 71 citations
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang et al.ICLR 2023 · 68 citations
Builds on2
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman et al.ICLR 2020 · 401 citations
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 112 citations
Related papers
- SLOG: A Structural Generalization Benchmark for Semantic ParsingBingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen et al.EMNLP 2023 · 3 citations
- Structural generalization in COGS: Supertagging is (almost) all you needAlban Petit, Caio F. Corro, François YvonEMNLP 2023
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
- Consistency Regularization Training for Compositional GeneralizationYongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng et al.ACL 2023 · 5 citations
- *-CFQ: Analyzing the Scalability of Machine Learning on a Compositional TaskDmitry Tsarkov, Tibor Tihon, Nathan Scales, Nikola Momchev et al.AAAI 2021 · 10 citations
