DP-Mix: Mixup-based Data Augmentation for Differentially Private Learning
Wenxuan Bao, Francesco Pittaluga, Vijay Kumar B. G, Vincent Bindschaedler
Abstract
Data augmentation techniques, such as simple image transformations and combinations, are highly effective at improving the generalization of computer vision models, especially when training data is limited. However, such techniques are fundamentally incompatible with differentially private learning approaches, due to the latter's built-in assumption that each training image's contribution to the learned model is bounded. In this paper, we investigate why naive applications of multi-sample data augmentation techniques, such as mixup, fail to achieve good performance and propose two novel data augmentation techniques specifically designed for the constraints of differentially private learning. Our first technique, DP-MIX SELF , achieves SoTA classification performance across a range of datasets and settings by performing mixup on self-augmented data. Our second technique, DP-MIX DIFF , further improves performance by incorporating synthetic data from a pre-trained diffusion model into the mixup process. We open-source the code at https://github.com/wenxuan-Bao/DP-Mix .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06f5d206-58ed-4d49-b94d-2e005ae3b0a8Cited by top-tier papers6
- Neural Collapse meets Differential Privacy: Curious behaviors of NoisyGD with Near-Perfect Representation LearningChendi Wang, Yuqing Zhu, Weijie J. Su, Yu-Xiang WangICML 2024 · 10 citations
- Understanding Private Learning From Feature PerspectiveMeng Ding, Mingxi Lei, Shaopeng Fu, Shaowei Wang et al.ICML 2026 · 2 citations
- Trustworthy Machine Learning through Data-Specific IndistinguishabilityHanshen Xiao, Zhen Yang, G. Edward SuhICML 2025
- Deep Learning with Plausible DeniabilityWenxuan Bao, Shan Jin, Hadi Abdullah, Anderson Nascimento et al.NeurIPS 2025
- Is Graph Mixup Beneficial? Investigating Interpolation And Empirical Performance of Graph Mixup MethodsSimon Forbat, Rainer GemullaICML 2026
Builds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- In-Distribution Public Data Synthesis With Diffusion Models for Differentially Private Image ClassificationJinseong Park, Yujin Choi, Jaewook LeeCVPR 2024
- dp-promise: Differentially Private Diffusion Probabilistic Models for Image SynthesisHaichen Wang, Shuchao Pang, Zhigang Lu, Yihang Rao et al.USENIX Security 2024 · 36 citations
- Enhance Image Classification via Inter-Class Image Mixup with Diffusion ModelZhicai Wang, Longhui Wei, Tan Wang, Heyu Chen et al.CVPR 2024 · 23 citations
- Effectively Using Public Data in Privacy Preserving Machine LearningMilad Nasr, Saeed Mahloujifar, Xinyu Tang, Prateek Mittal et al.ICML 2023 · 22 citations
- PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware PretrainingKecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao et al.USENIX Security 2024 · 23 citations
