Differentially Private Sharpness-Aware Training
Jinseong Park, Hoki Kim, Yujin Choi, Jaewook Lee
Abstract
Training deep learning models with differential privacy (DP) results in a degradation of performance. The training dynamics of models with DP show a significant difference from standard training, whereas understanding the geometric properties of private learning remains largely unexplored. In this paper, we investigate sharpness, a key factor in achieving better generalization, in private learning. We show that flat minima can help reduce the negative effects of per-example gradient clipping and the addition of Gaussian noise. We then verify the effectiveness of Sharpness-Aware Minimization (SAM) for seeking flat minima in private learning. However, we also discover that SAM is detrimental to the privacy budget and computational time due to its two-step optimization. Thus, we propose a new sharpness-aware training method that mitigates the privacy-optimization trade-off. Our experimental results demonstrate that the proposed method improves the performance of deep learning models with DP from both scratch and finetuning. Code is available at https://github. com/jinseongP/DPSAT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92c4731a-9a75-45b3-ade9-fb5d01385d2dCited by top-tier papers9
- Delving into Differentially Private TransformerYoulong Ding, Xueyang Wu, Yining Meng, Yonggang Luo et al.ICML 2024 · 11 citations
- Multi-Class Support Vector Machine with Differential PrivacyJinseong Park, Yujin Choi, Jaewook LeeNeurIPS 2025 · 1 citation
- DPQuant: Efficient and Private Model Training via Dynamic Quantization SchedulingYubo Gao, Renbo Tu, Gennady Pekhimenko, Nandita VijaykumarICLR 2026
- Sharpness-Aware Initialization: Improving Differentially Private Machine Learning from First PrinciplesZihao Wang, Rui Zhu, Dongruo Zhou, Zhikun Zhang et al.USENIX Security 2025
- RPGen: Robust and Differentially Private Synthetic Image GenerationZihao Wang, Hao Peng, Wei Dong, Yuecen Wei et al.AAAI 2026
Builds on31
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
Related papers
- Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization TermYun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao et al.KDD 2023 · 7 citations
- Stability Analysis of Sharpness-Aware MinimizationHoki Kim, Jinseong Park, Yujin Choi, Jaewook LeeICML 2026 · 18 citations
- Penalizing Gradient Norm for Efficiently Improving Generalization in Deep LearningYang Zhao, Hao Zhang, Xiuyuan HuICML 2022 · 165 citations
- Make Landscape Flatter in Differentially Private Federated LearningYifan Shi, Yingqi Liu, Kang Wei, Li Shen et al.CVPR 2023
- Enhancing DPSGD via Per-Sample Momentum and Low-Pass FilteringXincheng Xu, Thilina Ranbaduge, Qing Wang, Thierry Rakotoarivelo et al.AAAI 2026
