AutoDO: Robust AutoAugment for Biased Data With Label Noise via Scalable Probabilistic Implicit Differentiation
Denis A. Gudovskiy, Luca Rigazio, Shun Ishizaka, Kazuki Kozuka, Sotaro Tsukizawa
Abstract
AutoAugment [6] has sparked an interest in automated augmentation methods for deep learning models. These methods estimate image transformation policies for train data that improve generalization to test data. While recent papers evolved in the direction of decreasing policy search complexity, we show that those methods are not robust when applied to biased and noisy data. To overcome these limitations, we reformulate AutoAugment as a generalized automated dataset optimization (AutoDO) task that minimizes the distribution shift between test data and distorted train dataset. In our AutoDO model, we explicitly estimate a set of per-point hyperparameters to flexibly change distribution of train data. In particular, we include hyperparameters for augmentation, loss weights, and softlabels that are jointly estimated using implicit differentiation. We develop a theoretical probabilistic interpretation of this framework using Fisher information and show that its complexity scales linearly with the dataset size. Our experiments on SVHN, CIFAR-10/100, and ImageNet classification show up to 9.3% improvement for biased datasets with label noise compared to prior methods and, importantly, up to 36.6% gain for underrepresented SVHN classes 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- FedFixer: Mitigating Heterogeneous Label Noise in Federated LearningXinyuan Ji, Zhaowei Zhu, Wei Xi, Olga Gadyatskaya et al.AAAI 2024 · 30 citations
- Making Scalable Meta Learning PracticalSang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger et al.NeurIPS 2023 · 28 citations
- Hyperbolic Feature Augmentation via Distribution Estimation and Infinite Sampling on ManifoldsZhi Gao, Yuwei Wu, Yunde Jia, Mehrtash HarandiNeurIPS 2022 · 21 citations
- CUDA: Curriculum of Data Augmentation for Long-tailed RecognitionSumyeong Ahn, Jongwoo Ko, Se-Young YunICLR 2023 · 15 citations
- Efficient Scheduling of Data Augmentation for Deep Reinforcement LearningByungchan Ko, Jungseul OkNeurIPS 2022 · 6 citations
Builds on5
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- Online Hyper-Parameter Learning for Auto-Augmentation StrategyChen Lin, Minghao Guo, Chuming Li, Xin Yuan et al.ICCV 2019 · 92 citations
- Deep Active Learning for Biased Datasets via Fisher Kernel Self-SupervisionDenis A. Gudovskiy, Alec Hodgkinson, Takuya Yamaguchi, Sotaro TsukizawaCVPR 2020
Related papers
- AdaAug: Learning Class- and Instance-adaptive Data Augmentation PoliciesTsz-Him Cheung, Dit-Yan YeungICLR 2022 · 31 citations
- Adversarial AutoAugmentXinyu Zhang, Qiang Wang, Jian Zhang, Zhao ZhongICLR 2020 · 210 citations
- MetaAugment: Sample-Aware Data Augmentation Policy LearningFengwei Zhou, Jiawei Li, Chuanlong Xie, Fei Chen et al.AAAI 2021 · 35 citations
- SLACK: Stable Learning of Augmentations with Cold-Start and KL RegularizationJuliette Marrie, Michael Arbel, Diane Larlus, Julien MairalCVPR 2023
- AdaTransform: Adaptive Data TransformationZhiqiang Tang, Xi Peng, Tingfeng Li, Yizhe Zhu et al.ICCV 2019 · 20 citations
