Understanding Compositional Data Augmentation in Typologically Diverse Morphological Inflection
Farhan Samir, Miikka Silfverberg
Abstract
Data augmentation techniques are widely used in low-resource automatic morphological inflection to overcome data sparsity. However, the full implications of these techniques remain poorly understood. In this study, we aim to shed light on the theoretical aspects of the prominent data augmentation strategy STEM-CORRUPT (Silfverberg et al., 2017; Anastasopoulos and Neubig, 2019) , a method that generates synthetic examples by randomly substituting stem characters in gold standard training examples. To begin, we conduct an information-theoretic analysis, arguing that STEMCORRUPT improves compositional generalization by eliminating spurious correlations between morphemes, specifically between the stem and the affixes. Our theoretical analysis further leads us to study the sampleefficiency with which STEMCORRUPT reduces these spurious correlations. Through evaluation across seven typologically distinct languages, we demonstrate that selecting a subset of datapoints with both high diversity and high predictive uncertainty significantly enhances the data-efficiency of STEMCORRUPT. However, we also explore the impact of typological features on the choice of the data selection strategy and find that languages incorporating a high degree of allomorphy and phonological alternations derive less benefit from synthetic examples with high uncertainty. We attribute this effect to phonotactic violations induced by STEMCORRUPT, emphasizing the need for further research to ensure optimal performance across the entire spectrum of natural language morphology. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92e3ae78-18d8-4dba-84dc-674d92d447f0Cited by top-tier papers1
Ask how each one uses itBuilds on13
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 128 citations
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters et al.EMNLP 2021 · 72 citations
- Active Learning Helps Pretrained Models Learn the Intended TaskAlex Tamkin, Dat Nguyen, Salil Deshpande, Jesse Mu et al.NeurIPS 2022 · 54 citations
- Learning to Recombine and Resample Data For Compositional GeneralizationEkin Akyürek, Afra Feyza Akyürek, Jacob AndreasICLR 2021 · 36 citations
Related papers
- Is linguistically-motivated data augmentation worth it?Ray Groshan, Michael Ginn, Alexis PalmerACL 2025
- Bootstrapping Techniques for Polysynthetic Morphological AnalysisWilliam Lane, Steven BirdACL 2020
- Improving Low-Resource Morphological Inflection via Self-Supervised ObjectivesAdam Wiemerslage, Katharina von der WenseACL 2025 · 2 citations
- Modeling Morphological Typology for Unsupervised Learning of Language MorphologyHongzhi Xu, Jordan Kodner, Mitchell Marcus, Charles YangACL 2020 · 7 citations
- Morphological Inflection: A Reality CheckJordan Kodner, Sarah R. B. Payne, Salam Khalifa, Zoey LiuACL 2023 · 6 citations
