What Makes Better Augmentation Strategies? Augment Difficult but Not too Different
Jaehyung Kim, Dongyeop Kang, Sungsoo Ahn, Jinwoo Shin
Abstract
The practice of data augmentation has been extensively used to boost the performance of deep neural networks for various NLP tasks. It is more effective when only a limited number of labeled samples is available, e.g., low-data or class-imbalanced regimes. Most current augmentation techniques rely on parameter tuning or inherent randomness; hence, their effectiveness largely varies on the tasks. To efficiently find the best augmentation strategy for each task, learning data augmentation policy is a promising solution, but the question of what makes a good augmentation in NLP tasks and how to design the reward function for learning a good policy remains under-explored. To answer this, we hypothesize that good data augmentation should construct more diverse and challenging samples for providing informative training signals, while avoiding the risk of losing the semantics of original samples. Therefore, we design a novel reward function for updating the augmentation policy to construct difficult but not too different samples (DND). Particularly, we jointly optimize a data augmentation policy while training the model, to construct the augmented samples with low confidence but a high semantic similarity with original ones. In addition, we introduce a sample re-weighting scheme to focus on difficult augmented samples after the original ones are learned confidently for more effective learning from the augmented ones. Our learning-based augmentation outperforms the recent state-of-the-art augmentation schemes on various text classification tasks and GLUE benchmark by successfully discovering the effective augmentations for each task. Remarkably, our method is more effective on the challenging low-data and class-imbalanced regimes, and the learned augmentation policy is well-transferable to the different tasks and models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b6c6c3da-ce6d-4ddc-b771-55635bb9395dCited by top-tier papers3
- Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement LearningGuozheng Ma, Linrui Zhang, Haoyu Wang, Lu Li et al.NeurIPS 2023 · 24 citations
- How Much Data Are Augmentations Worth? An Investigation into Scaling Laws, Invariance, and Implicit RegularizationJonas Geiping, Micah Goldblum, Gowthami Somepalli, Ravid Shwartz-Ziv et al.ICLR 2023 · 11 citations
- Symmetric Replay Training: Enhancing Sample Efficiency in Deep Reinforcement Learning for Combinatorial OptimizationHyeonah Kim, Minsu Kim, Sungsoo Ahn, Jinkyoo ParkICML 2024 · 9 citations
Related papers
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationRuibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma et al.EMNLP 2020 · 62 citations
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 3 citations
- Text AutoAugment: Learning Compositional Augmentation Policy for Text ClassificationShuhuai Ren, Jinchao Zhang, Lei Li, Xu Sun et al.EMNLP 2021 · 22 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and BeyondZhengjie Miao, Yuliang Li, Xiaolan WangSIGMOD 2021 · 63 citations
