Learning to Generate Image Embeddings with User-Level Differential Privacy
Zheng Xu, Maxwell D. Collins, Yuxiao Wang, Liviu Panait, Sewoong Oh, Sean Augenstein, Ting Liu, Florian Schroff, H. Brendan McMahan
摘要
Small on-device models have been successfully trained with user-level differential privacy (DP) for next word prediction and image classification tasks in the past. However, existing methods can fail when directly applied to learn embedding models using supervised training data with a large class space. To achieve user-level DP for large imageto-embedding feature extractors, we propose DP-FedEmb, a variant of federated learning algorithms with per-user sensitivity control and noise addition, to train from userpartitioned data centralized in the datacenter. DP-FedEmb combines virtual clients, partial aggregation, private local fine-tuning, and public pretraining to achieve strong privacy utility trade-offs. We apply DP-FedEmb to train image embedding models for faces, landmarks and natural species, and demonstrate its superior utility under same privacy budget on benchmark datasets DigiFace, EMNIST, GLD and iNaturalist. We further illustrate it is possible to achieve strong user-level DP guarantees of ϵ < 2 while controlling the utility drop within 5%, when millions of users can participate in training .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Privacy-Preserving Instructions for Aligning Large Language ModelsDa Yu, Peter Kairouz, Sewoong Oh, Zheng XuICML 2024 · 被引用 41 次
- PrE-Text: Training Language Models on Private Federated Data in the Age of LLMsCharlie Hou, Akshat Shrivastava, Hongyuan Zhan, Rylan Conway 等ICML 2024 · 被引用 30 次
- ClassID: Enabling Student Behavior Attribution from Ambient Classroom Sensing SystemsPrasoon Patidar, Tricia J. Ngoon, John Zimmerman, Amy Ogan 等UbiComp 2024 · 被引用 7 次
- Share Your Representation Only: Guaranteed Improvement of the Privacy-Utility Tradeoff in Federated LearningZebang Shen, Jiayuan Ye, Anmin Kang, Hamed Hassani 等ICLR 2023 · 被引用 6 次
- Differentially Private 2D Human Pose EstimationKaushik Bhargav Sivangi, Paul Henderson, Fani DeligianniCVPR 2026 · 被引用 1 次
它引用的顶会 Paper32
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
相关 Paper
- Sparsity-Preserving Differentially Private Training of Large Embedding ModelsBadih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar 等NeurIPS 2023 · 被引用 9 次
- Sanitizing Sentence Embeddings (and Labels) for Local Differential PrivacyMinxin Du, Xiang Yue, Sherman S. M. Chow, Huan SunWWW 2023 · 被引用 26 次
- Private Model Personalization RevisitedConor Snedeker, Xinyu Zhou, Raef BassilyICML 2025
- Sentence-level Privacy for Document EmbeddingsCasey Meehan, Khalil Mrini, Kamalika ChaudhuriACL 2022 · 被引用 26 次
- Do not Let Privacy Overbill Utility: Gradient Embedding Perturbation for Private LearningDa Yu, Huishuai Zhang, Wei Chen, Tie-Yan LiuICLR 2021 · 被引用 133 次
