Learning to Generate Image Embeddings with User-Level Differential Privacy
Zheng Xu, Maxwell D. Collins, Yuxiao Wang, Liviu Panait, Sewoong Oh, Sean Augenstein, Ting Liu, Florian Schroff, H. Brendan McMahan
Abstract
Small on-device models have been successfully trained with user-level differential privacy (DP) for next word prediction and image classification tasks in the past. However, existing methods can fail when directly applied to learn embedding models using supervised training data with a large class space. To achieve user-level DP for large imageto-embedding feature extractors, we propose DP-FedEmb, a variant of federated learning algorithms with per-user sensitivity control and noise addition, to train from userpartitioned data centralized in the datacenter. DP-FedEmb combines virtual clients, partial aggregation, private local fine-tuning, and public pretraining to achieve strong privacy utility trade-offs. We apply DP-FedEmb to train image embedding models for faces, landmarks and natural species, and demonstrate its superior utility under same privacy budget on benchmark datasets DigiFace, EMNIST, GLD and iNaturalist. We further illustrate it is possible to achieve strong user-level DP guarantees of ϵ < 2 while controlling the utility drop within 5%, when millions of users can participate in training .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Privacy-Preserving Instructions for Aligning Large Language ModelsDa Yu, Peter Kairouz, Sewoong Oh, Zheng XuICML 2024 · 41 citations
- PrE-Text: Training Language Models on Private Federated Data in the Age of LLMsCharlie Hou, Akshat Shrivastava, Hongyuan Zhan, Rylan Conway et al.ICML 2024 · 30 citations
- ClassID: Enabling Student Behavior Attribution from Ambient Classroom Sensing SystemsPrasoon Patidar, Tricia J. Ngoon, John Zimmerman, Amy Ogan et al.UbiComp 2024 · 7 citations
- Share Your Representation Only: Guaranteed Improvement of the Privacy-Utility Tradeoff in Federated LearningZebang Shen, Jiayuan Ye, Anmin Kang, Hamed Hassani et al.ICLR 2023 · 6 citations
- Differentially Private 2D Human Pose EstimationKaushik Bhargav Sivangi, Paul Henderson, Fani DeligianniCVPR 2026 · 1 citation
Builds on32
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
Related papers
- Sparsity-Preserving Differentially Private Training of Large Embedding ModelsBadih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar et al.NeurIPS 2023 · 9 citations
- Sanitizing Sentence Embeddings (and Labels) for Local Differential PrivacyMinxin Du, Xiang Yue, Sherman S. M. Chow, Huan SunWWW 2023 · 26 citations
- Private Model Personalization RevisitedConor Snedeker, Xinyu Zhou, Raef BassilyICML 2025
- Sentence-level Privacy for Document EmbeddingsCasey Meehan, Khalil Mrini, Kamalika ChaudhuriACL 2022 · 26 citations
- Do not Let Privacy Overbill Utility: Gradient Embedding Perturbation for Private LearningDa Yu, Huishuai Zhang, Wei Chen, Tie-Yan LiuICLR 2021 · 133 citations
