Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
Mingyang Yi, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, Zhi-Ming Ma
摘要
Data augmentation is an effective technique to improve the generalization of deep neural networks. However, previous data augmentation methods usually treat the augmented samples equally without considering their individual impacts on the model. To address this, for the augmented samples from the same training example, we propose to assign different weights to them. We construct the maximal expected loss which is the supremum over any reweighted loss on augmented samples. Inspired by adversarial training, we minimize this maximal expected loss (MMEL) and obtain a simple and interpretable closed-form solution: more attention should be paid to augmented samples with large loss values (i.e., harder examples). Minimizing this maximal expected loss enables the model to perform well under any reweighting strategy. The proposed method can generally be applied on top of any data augmentation methods. Experiments are conducted on both natural language understanding tasks with token-level data augmentation, and image classification tasks with commonly-used image augmentation techniques like random crop and horizontal flip. Empirical results show that the proposed method improves the generalization performance of the model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Doubly Robust Instance-Reweighted Adversarial TrainingDaouda Sow, Sen Lin, Zhangyang Wang, Yingbin LiangICLR 2024 · 被引用 2 次
- Dynamic Loss-Based Sample Reweighting for Improved Large Language Model PretrainingDaouda Sow, Herbert Woisetschläger, Saikiran Bulusu, Shiqiang Wang 等ICLR 2025
- Soft Augmentation for Image ClassificationYang Liu, Shen Yan, Laura Leal-Taixé, James Hays 等CVPR 2023
- Breaking Correlation Shift via Conditional Invariant RegularizerMingyang Yi, Ruoyu Wang, Jiacheng Sun, Zhenguo Li 等ICLR 2023
它引用的顶会 Paper5
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun 等ICLR 2020 · 被引用 502 次
- Beyond Synthetic Noise: Deep Learning on Controlled Noisy LabelsLu Jiang, Di Huang, Mason Liu, Weilong YangICML 2020 · 被引用 241 次
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu 等ACL 2020 · 被引用 148 次
- AdvAug: Robust Adversarial Augmentation for Neural Machine TranslationYong Cheng, Lu Jiang, Wolfgang Macherey, Jacob EisensteinACL 2020 · 被引用 105 次
相关 Paper
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- MaxUp: Lightweight Adversarial Training With Data Augmentation Improves Neural Network TrainingChengyue Gong, Tongzheng Ren, Mao Ye, Qiang LiuCVPR 2021
- TeachAugment: Data Augmentation Optimization Using Teacher KnowledgeTeppei SuzukiCVPR 2022 · 被引用 57 次
- M2m: Imbalanced Classification via Major-to-Minor TranslationJaehyung Kim, Jongheon Jeong, Jinwoo ShinCVPR 2020
- MetaAugment: Sample-Aware Data Augmentation Policy LearningFengwei Zhou, Jiawei Li, Chuanlong Xie, Fei Chen 等AAAI 2021 · 被引用 35 次
