Data Factors for Better Compositional Generalization
Xiang Zhou, Yichen Jiang, Mohit Bansal
Abstract
Recent diagnostic datasets on compositional generalization, such as SCAN (Lake and Baroni, 2018) and COGS (Kim and Linzen, 2020), expose severe problems in models trained from scratch on these datasets. However, in contrast to this poor performance, state-of-the-art models trained on larger and more general datasets show better generalization ability. In this work, to reconcile this inconsistency, we conduct an empirical analysis by training Transformer models on a variety of training sets with different data factors, including dataset scale, pattern complexity, example difficulty, etc. First, we show that increased dataset complexity can lead to better generalization behavior on multiple different generalization challenges. To further understand this improvement, we show two axes of the benefit from more complex datasets: they provide more diverse examples so compositional understanding becomes more effective, and they also prevent ungeneralizable memorization of the examples due to reduced example repetition frequency. Finally, we explore how training examples of different difficulty levels influence generalization differently. On synthetic datasets, simple examples invoke stronger compositionality than hard examples do. On larger-scale real language datasets, while hard examples become more important potentially to ensure decent data coverage, a balanced mixture of simple and hard examples manages to induce the strongest generalizability. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f21d868-57a8-439b-a4f2-6d24a9c6e41cCited by top-tier papers6
- The Unreasonable Effectiveness of Easy Training Data for Hard TasksPeter Hase, Mohit Bansal, Peter Clark, Sarah WiegreffeACL 2024 · 3 citations
- Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMsEva VanmassenhoveACL 2026 · 1 citation
- Data Distributional Properties As Inductive Bias for Systematic GeneralizationFelipe del Río, Alain Raymond-Saez, Daniel Florea, Rodrigo Toro Icarte et al.CVPR 2025
- Does Data Scaling Lead to Visual Compositional Generalization?Arnas Uselis, Andrea Dittadi, Seong Joon OhICML 2025
- CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information AlignmentNura Aljaafari, Danilo S. Carvalho, André FreitasEMNLP 2025
Builds on26
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang et al.NeurIPS 2022 · 407 citations
Related papers
- Mutual Exclusivity Training and Primitive Augmentation to Induce CompositionalityYichen Jiang, Xiang Zhou, Mohit BansalEMNLP 2022 · 1 citation
- *-CFQ: Analyzing the Scalability of Machine Learning on a Compositional TaskDmitry Tsarkov, Tibor Tihon, Nathan Scales, Nikola Momchev et al.AAAI 2021 · 10 citations
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of TransformersRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberEMNLP 2021 · 55 citations
- The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case StudyVerna Dankers, Elia Bruni, Dieuwke HupkesACL 2022
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
