CHEF in the Language Kitchen: A Generative Data Augmentation Leveraging Korean Morpheme Ingredients
Jaehyung Seo, Hyeonseok Moon, Jaewook Lee, Sugyeong Eo, Chanjun Park, Heuiseok Lim
Abstract
Korean morphological variations present unique opportunities and challenges in natural language processing (NLP), necessitating an advanced understanding of morpheme-based sentence construction. The complexity of morphological variations allows for diverse sentence forms based on the syntactic-semantic integration of functional morphemes (i.e., affixes) to lexical morphemes (i.e., roots). With this in mind, we propose a method - CHEF, replicating the morphological transformations inherent in sentences based on lexical and functional morpheme combinations through generative data augmentation. CHEF operates using a morpheme blender and a label discriminator, thereby enhancing the diversity of Korean sentence forms by capturing the properties of agglutination while maintaining label consistency. We conduct experiments on Korean multiple classification datasets, improving model performance in full- and few-shot settings. Our proposed method boosts performance beyond the preceding data augmentation methods without incurring external data usage. We demonstrate that our approach achieves comparable results yielded by augmentation techniques that use large language models (LLMs).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in KoreanSungHo Kim, Nayeon Kim, Taehee Jeon, SangKeun LeeACL 2025
- Is linguistically-motivated data augmentation worth it?Ray Groshan, Michael Ginn, Alexis PalmerACL 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Nonlinear Mixup: Out-Of-Manifold Data Augmentation for Text ClassificationHongyu GuoAAAI 2020 · 124 citations
Related papers
- ALP: Data Augmentation Using Lexicalized PCFGs for Few-Shot Text ClassificationHazel H. Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha et al.AAAI 2022 · 33 citations
- DiNeR: A Large Realistic Dataset for Evaluating Compositional GeneralizationChengang Hu, Xiao Liu, Yansong FengEMNLP 2023
- Enhance Image Classification via Inter-Class Image Mixup with Diffusion ModelZhicai Wang, Longhui Wei, Tan Wang, Heyu Chen et al.CVPR 2024 · 23 citations
- Learning to Recombine and Resample Data For Compositional GeneralizationEkin Akyürek, Afra Feyza Akyürek, Jacob AndreasICLR 2021 · 36 citations
- PromptMix: A Class Boundary Augmentation Method for Large Language Model DistillationGaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. LaradjiEMNLP 2023 · 11 citations
