Mutual Exclusivity Training and Primitive Augmentation to Induce Compositionality
Yichen Jiang, Xiang Zhou, Mohit Bansal
摘要
Recent datasets expose the lack of the systematic generalization ability in standard sequence-to-sequence models. In this work, we analyze this behavior of seq2seq models and identify two contributing factors: a lack of mutual exclusivity bias (one target sequence can only be mapped to one source sequence), and the tendency to memorize whole examples rather than separating structures from contents. We propose two techniques to address these two issues respectively: Mutual Exclusivity Training that prevents the model from producing seen generations when facing novel examples via an unlikelihood-based loss, and prim2primX data augmentation that automatically diversifies the arguments of every syntactic function to prevent memorizing and provide a compositional inductive bias without exposing test-set data. Combining these two techniques, we show substantial empirical improvements using standard sequence-to-sequence models (LSTMs and Transformers) on two widely-used compositionality datasets: SCAN and COGS. Finally, we provide analysis characterizing the improvements as well as the remaining challenges, and provide detailed ablations of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning to Substitute Spans towards Improving Compositional GeneralizationZhaoyi Li, Ying Wei, Defu LianACL 2023 · 被引用 3 次
- Data Factors for Better Compositional GeneralizationXiang Zhou, Yichen Jiang, Mohit BansalEMNLP 2023 · 被引用 2 次
- Data Distributional Properties As Inductive Bias for Systematic GeneralizationFelipe del Río, Alain Raymond-Saez, Daniel Florea, Rodrigo Toro Icarte 等CVPR 2025
- Inducing Systematicity in Transformers by Attending to Structurally Quantized EmbeddingsYichen Jiang, Xiang Zhou, Mohit BansalACL 2024
它引用的顶会 Paper17
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- Learning Compositional Rules via Neural Program SynthesisMaxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. LakeNeurIPS 2020 · 被引用 120 次
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 被引用 112 次
相关 Paper
- Good-Enough Compositional Data AugmentationJacob AndreasACL 2020 · 被引用 15 次
- What they do when in doubt: a study of inductive biases in seq2seq learnersEugene Kharitonov, Rahma ChaabouniICLR 2021 · 被引用 29 次
- Learning to Recombine and Resample Data For Compositional GeneralizationEkin Akyürek, Afra Feyza Akyürek, Jacob AndreasICLR 2021 · 被引用 36 次
- Meta-Learning to Compositionally GeneralizeHenry Conklin, Bailin Wang, Kenny Smith, Ivan TitovACL 2021
- Inducing Transformer's Compositional Generalization Ability via Auxiliary Sequence Prediction TasksYichen Jiang, Mohit BansalEMNLP 2021
