COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
Najoung Kim, Tal Linzen
摘要
Natural language is characterized by compositionality: the meaning of a complex expression is constructed from the meanings of its constituent parts. To facilitate the evaluation of the compositional abilities of language processing architectures, we introduce COGS, a semantic parsing dataset based on a fragment of English. The evaluation portion of COGS contains multiple systematic gaps that can only be addressed by compositional generalization; these include new combinations of familiar syntactic structures, or new combinations of familiar words and familiar structures. In experiments with Transformers and LSTMs, we found that in-distribution accuracy on the COGS test set was near-perfect (96-99%), but generalization accuracy was substantially lower (16-35%) and showed high sensitivity to random seed (±6-8%). These findings indicate that contemporary standard NLP models are limited in their compositional generalization capacity, and position COGS as a good way to measure progress.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper96
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao 等ICLR 2022 · 被引用 237 次
- Making Transformers Solve Compositional TasksSantiago Ontañón, Joshua Ainslie, Zachary Fisher, Vaclav CvicekACL 2022 · 被引用 87 次
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang 等NeurIPS 2025 · 被引用 71 次
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang 等ICLR 2023 · 被引用 68 次
它引用的顶会 Paper2
- Measuring Compositional Generalization: A Comprehensive Method on Realistic DataDaniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman 等ICLR 2020 · 被引用 401 次
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 被引用 112 次
相关 Paper
- SLOG: A Structural Generalization Benchmark for Semantic ParsingBingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen 等EMNLP 2023 · 被引用 3 次
- Structural generalization in COGS: Supertagging is (almost) all you needAlban Petit, Caio F. Corro, François YvonEMNLP 2023
- Compositional Generalization without Trees using Multiset Tagging and Latent PermutationsMatthias Lindemann, Alexander Koller, Ivan TitovACL 2023
- Consistency Regularization Training for Compositional GeneralizationYongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng 等ACL 2023 · 被引用 5 次
- *-CFQ: Analyzing the Scalability of Machine Learning on a Compositional TaskDmitry Tsarkov, Tibor Tihon, Nathan Scales, Nikola Momchev 等AAAI 2021 · 被引用 10 次
