Effective Data Augmentation With Diffusion Models
Brandon Trabucco, Kyle Doherty, Max Gurinas, Ruslan Salakhutdinov
Abstract
Data augmentation is one of the most prevalent tools in deep learning, underpinning many recent advances, including those from classification, generative models, and representation learning. The standard approach to data augmentation combines simple transformations like rotations and flips to generate new images from existing ones. However, these new images lack diversity along key semantic axes present in the data. Current augmentations cannot alter the high-level semantic attributes, such as animal species present in a scene, to enhance the diversity of data. We address the lack of diversity in data augmentation with image-to-image transformations parameterized by pre-trained text-to-image diffusion models. Our method edits images to change their semantics using an off-the-shelf diffusion model, and generalizes to novel visual concepts from a few labelled examples. We evaluate our approach on few-shot image classification tasks, and on a real-world weed recognition task, and observe an improvement in accuracy in tested domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers98
- Global Structure-Aware Diffusion Process for Low-light Image EnhancementJinhui Hou, Zhiyu Zhu, Junhui Hou, Hui Liu et al.NeurIPS 2023 · 280 citations
- Diversify Your Vision Datasets with Automatic Diffusion-based AugmentationLisa Dunlap, Alyssa Umino, Han Zhang, Jiezhi Yang et al.NeurIPS 2023 · 136 citations
- Understanding Hallucinations in Diffusion Models through Mode InterpolationSumukh K. Aithal, Pratyush Maini, Zachary C. Lipton, J. Zico KolterNeurIPS 2024 · 121 citations
- Dream the Impossible: Outlier Imagination with Diffusion ModelsXuefeng Du, Yiyou Sun, Jerry Zhu, Yixuan LiNeurIPS 2023 · 114 citations
- FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation ModelsLihe Yang, Xiaogang Xu, Bingyi Kang, Yinghuan Shi et al.NeurIPS 2023 · 94 citations
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Enhance Image Classification via Inter-Class Image Mixup with Diffusion ModelZhicai Wang, Longhui Wei, Tan Wang, Heyu Chen et al.CVPR 2024 · 23 citations
- OntoAug: Rethinking Generative Data Augmentation via Ontology GuidanceShuo Wang, Zhichuan Wang, Jun LuoCVPR 2026
- Generating Images of Rare Concepts Using Pre-trained Diffusion ModelsDvir Samuel, Rami Ben-Ari, Simon Raviv, Nir Darshan et al.AAAI 2024 · 82 citations
- Plug-and-Play Diffusion Features for Text-Driven Image-to-Image TranslationNarek Tumanyan, Michal Geyer, Shai Bagon, Tali DekelCVPR 2023
- DRMix: Decomposition-Recomposition Data Augmentation with Diffusion ModelShuo Wang, Zhichuan Wang, Yanmin Chen, Mengyao Zhou et al.ACM MM 2025
