Understanding Compositional Data Augmentation in Typologically Diverse Morphological Inflection
Farhan Samir, Miikka Silfverberg
摘要
Data augmentation techniques are widely used in low-resource automatic morphological inflection to overcome data sparsity. However, the full implications of these techniques remain poorly understood. In this study, we aim to shed light on the theoretical aspects of the prominent data augmentation strategy STEM-CORRUPT (Silfverberg et al., 2017; Anastasopoulos and Neubig, 2019) , a method that generates synthetic examples by randomly substituting stem characters in gold standard training examples. To begin, we conduct an information-theoretic analysis, arguing that STEMCORRUPT improves compositional generalization by eliminating spurious correlations between morphemes, specifically between the stem and the affixes. Our theoretical analysis further leads us to study the sampleefficiency with which STEMCORRUPT reduces these spurious correlations. Through evaluation across seven typologically distinct languages, we demonstrate that selecting a subset of datapoints with both high diversity and high predictive uncertainty significantly enhances the data-efficiency of STEMCORRUPT. However, we also explore the impact of typological features on the choice of the data selection strategy and find that languages incorporating a high degree of allomorphy and phonological alternations derive less benefit from synthetic examples with high uncertainty. We attribute this effect to phonotactic violations induced by STEMCORRUPT, emphasizing the need for further research to ensure optimal performance across the entire spectrum of natural language morphology. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters 等EMNLP 2021 · 被引用 72 次
- Active Learning Helps Pretrained Models Learn the Intended TaskAlex Tamkin, Dat Nguyen, Salil Deshpande, Jesse Mu 等NeurIPS 2022 · 被引用 54 次
- Learning to Recombine and Resample Data For Compositional GeneralizationEkin Akyürek, Afra Feyza Akyürek, Jacob AndreasICLR 2021 · 被引用 36 次
相关 Paper
- Is linguistically-motivated data augmentation worth it?Ray Groshan, Michael Ginn, Alexis PalmerACL 2025
- Bootstrapping Techniques for Polysynthetic Morphological AnalysisWilliam Lane, Steven BirdACL 2020
- Improving Low-Resource Morphological Inflection via Self-Supervised ObjectivesAdam Wiemerslage, Katharina von der WenseACL 2025 · 被引用 2 次
- Modeling Morphological Typology for Unsupervised Learning of Language MorphologyHongzhi Xu, Jordan Kodner, Mitchell Marcus, Charles YangACL 2020 · 被引用 7 次
- Morphological Inflection: A Reality CheckJordan Kodner, Sarah R. B. Payne, Salam Khalifa, Zoey LiuACL 2023 · 被引用 6 次
