DataLens: Scalable Privacy Preserving Training via Gradient Compression and Aggregation
Boxin Wang, Fan Wu, Yunhui Long, Luka Rimanic, Ce Zhang, Bo Li
摘要
Recent success of deep neural networks (DNNs) hinges on the availability of large-scale dataset; however, training on such dataset often poses privacy risks for sensitive training information. In this paper, we aim to explore the power of generative models and gradient sparsity, and propose a scalable privacy-preserving generative model DataLens, which is able to generate synthetic data in a differentially private (DP) way given sensitive input data. Thus, it is possible to train models for different down-stream tasks with the generated data while protecting the private information. In particular, we leverage the generative adversarial networks (GAN) and PATE framework to train multiple discriminators as "teacher" models, allowing them to vote with their gradient vectors to guarantee privacy. Comparing with the standard PATE privacy preserving framework which allows teachers to vote on one-dimensional predictions, voting on the high dimensional gradient vectors is challenging in terms of privacy preservation. As dimension reduction techniques are required, we need to navigate a delicate tradeoff space between (1) the improvement of privacy preservation and (2) the slowdown of SGD convergence. To tackle this, we propose a novel dimension compression and aggregation approach TopAgg, which combines top-k dimension compression with a corresponding noise injection mechanism. We theoretically prove that the DataLens framework guarantees differential privacy for its generated data, and provide a novel analysis on its convergence to illustrate such a tradeoff on privacy and convergence rate, which requires non-trivial analysis as it requires a joint analysis on gradient compression, coordinate-wise gradient clipping, and DP mechanism. To demonstrate the practical usage of DataLens, we conduct extensive experiments on diverse datasets including MNIST, Fashion-MNIST, and high dimensional CelebA and Place365 datasets. We show that DataLens significantly outperforms other baseline differentially private data generative models. Our code is publicly available at https://github.com/AI-secure/DataLens.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn DivergenceTianshi Cao, Alex Bie, Arash Vahdat, Sanja Fidler 等NeurIPS 2021 · 被引用 88 次
- SoK: Privacy-Preserving Data SynthesisYuzheng Hu, Fan Wu, Qinbin Li, Yunhui Long 等S&P 2024 · 被引用 61 次
- Private Set Generation with Discriminative InformationDingfan Chen, Raouf Kerkouche, Mario FritzNeurIPS 2022 · 被引用 51 次
- dp-promise: Differentially Private Diffusion Probabilistic Models for Image SynthesisHaichen Wang, Shuchao Pang, Zhigang Lu, Yihang Rao 等USENIX Security 2024 · 被引用 36 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
它引用的顶会 Paper13
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Differentially Private Learning with Adaptive ClippingGalen Andrew, Om Thakkar, Brendan McMahan, Swaroop RamaswamyNeurIPS 2021 · 被引用 425 次
- FetchSGD: Communication-Efficient Federated Learning with SketchingDaniel Rothchild, Ashwinee Panda, Enayat Ullah, Nikita Ivkin 等ICML 2020 · 被引用 425 次
相关 Paper
- G-PATE: Scalable Differentially Private Data Generator via Private Aggregation of Teacher DiscriminatorsYunhui Long, Boxin Wang, Zhuolin Yang, Bhavya Kailkhura 等NeurIPS 2021 · 被引用 91 次
- PEARL: Data Synthesis via Private Embeddings and Adversarial Reconstruction LearningSeng Pei Liew, Tsubasa Takahashi, Michihiko UenoICLR 2022 · 被引用 32 次
- PateGAIL++: Utility Optimized Private Trajectory Generation with Imitation LearningYingjie Ma, Bijal Bharadva, Xin Zhang, Joann Qiongna ChenICLR 2026
- In Differential Privacy, There is Truth: on Vote-Histogram Leakage in Ensemble Private LearningJiaqi Wang, Roei Schuster, Ilia Shumailov, David Lie 等NeurIPS 2022 · 被引用 8 次
- SeqPATE: Differentially Private Text Generation via Knowledge DistillationZhiliang Tian, Yingxiu Zhao, Ziyue Huang, Yu-Xiang Wang 等NeurIPS 2022 · 被引用 29 次
