Data Augmentation for Text Generation Without Any Augmented Data
Wei Bi, Huayang Li, Jiacheng Huang
Abstract
Data augmentation is an effective way to improve the performance of many neural text generation models. However, current data augmentation methods need to define or choose proper data mapping functions that map the original samples into the augmented samples. In this work, we derive an objective to formulate the problem of data augmentation on text generation tasks without any use of augmented data constructed by specific mapping functions. Our proposed objective can be efficiently optimized and applied to popular loss functions on text generation tasks with a convergence rate guarantee. Experiments on five datasets of two text generation tasks show that our approach can approximate or even surpass popular data augmentation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53800ad6-dc28-4e88-93a3-6cb3fe1dd9ceCited by top-tier papers1
Ask how each one uses itBuilds on2
- Data Manipulation: Towards Effective Instance Learning for Neural Dialogue Generation via Learning to Augment and ReweightHengyi Cai, Hongshen Chen, Yonghao Song, Cheng Zhang et al.ACL 2020 · 57 citations
- Dialogue Distillation: Open-Domain Dialogue Augmentation Using Unpaired DataRongsheng Zhang, Yinhe Zheng, Jianzhi Shao, Xiaoxi Mao et al.EMNLP 2020 · 25 citations
Related papers
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian et al.ACL 2022 · 91 citations
- Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text ClassificationHsun-Yu Kuo, Yin-Hsiang Liao, Yu-Chieh Chao, Wei-Yun Ma et al.ICLR 2025
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationRuibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma et al.EMNLP 2020 · 62 citations
- Beyond MLE: Convex Learning for Text GenerationChenze Shao, Zhengrui Ma, Min Zhang, Yang FengNeurIPS 2023 · 5 citations
