Mutual Exclusivity Training and Primitive Augmentation to Induce Compositionality
Yichen Jiang, Xiang Zhou, Mohit Bansal
Abstract
Recent datasets expose the lack of the systematic generalization ability in standard sequence-to-sequence models. In this work, we analyze this behavior of seq2seq models and identify two contributing factors: a lack of mutual exclusivity bias (one target sequence can only be mapped to one source sequence), and the tendency to memorize whole examples rather than separating structures from contents. We propose two techniques to address these two issues respectively: Mutual Exclusivity Training that prevents the model from producing seen generations when facing novel examples via an unlikelihood-based loss, and prim2primX data augmentation that automatically diversifies the arguments of every syntactic function to prevent memorizing and provide a compositional inductive bias without exposing test-set data. Combining these two techniques, we show substantial empirical improvements using standard sequence-to-sequence models (LSTMs and Transformers) on two widely-used compositionality datasets: SCAN and COGS. Finally, we provide analysis characterizing the improvements as well as the remaining challenges, and provide detailed ablations of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fdcd509-84b1-4010-997d-3d293a6d8a0bCited by top-tier papers4
- Learning to Substitute Spans towards Improving Compositional GeneralizationZhaoyi Li, Ying Wei, Defu LianACL 2023 · 3 citations
- Data Factors for Better Compositional GeneralizationXiang Zhou, Yichen Jiang, Mohit BansalEMNLP 2023 · 2 citations
- Data Distributional Properties As Inductive Bias for Systematic GeneralizationFelipe del Río, Alain Raymond-Saez, Daniel Florea, Rodrigo Toro Icarte et al.CVPR 2025
- Inducing Systematicity in Transformers by Attending to Structurally Quantized EmbeddingsYichen Jiang, Xiang Zhou, Mohit BansalACL 2024
Builds on17
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Learning Compositional Rules via Neural Program SynthesisMaxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. LakeNeurIPS 2020 · 120 citations
- Permutation Equivariant Models for Compositional Generalization in LanguageJonathan Gordon, David Lopez-Paz, Marco Baroni, Diane BouchacourtICLR 2020 · 112 citations
Related papers
- Good-Enough Compositional Data AugmentationJacob AndreasACL 2020 · 15 citations
- What they do when in doubt: a study of inductive biases in seq2seq learnersEugene Kharitonov, Rahma ChaabouniICLR 2021 · 29 citations
- Learning to Recombine and Resample Data For Compositional GeneralizationEkin Akyürek, Afra Feyza Akyürek, Jacob AndreasICLR 2021 · 36 citations
- Meta-Learning to Compositionally GeneralizeHenry Conklin, Bailin Wang, Kenny Smith, Ivan TitovACL 2021
- Inducing Transformer's Compositional Generalization Ability via Auxiliary Sequence Prediction TasksYichen Jiang, Mohit BansalEMNLP 2021
