InstaHide: Instance-hiding Schemes for Private Distributed Learning
Yangsibo Huang, Zhao Song, Kai Li, Sanjeev Arora
Abstract
How can multiple distributed entities collaboratively train a shared deep net on their private data while preserving privacy? This paper introduces InstaHide, a simple encryption of training images, which can be plugged into existing distributed deep learning pipelines. The encryption is efficient and applying it during training has minor effect on test accuracy. InstaHide encrypts each training image with a "one-time secret key" which consists of mixing a number of randomly chosen images and applying a random pixel-wise mask. Other contributions of this paper include: (a) Using a large public dataset (e.g. ImageNet) for mixing during its encryption, which improves security. (b) Experimental results to show effectiveness in preserving privacy against known attacks with only minor effects on accuracy. (c) Theoretical analysis showing that successfully attacking privacy requires attackers to solve a difficult computational problem. (d) Demonstrating that use of the pixel-wise mask is important for security, since Mixup alone is shown to be insecure to some some efficient attacks. (e) Release of a challenge dataset 1 to encourage new attacks. Our code is available at https://github.com/Hazelsuko07/InstaHide .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers38
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Evaluating Gradient Inversion Attacks and Defenses in Federated LearningYangsibo Huang, Samyak Gupta, Zhao Song, Kai Li et al.NeurIPS 2021 · 419 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir et al.NeurIPS 2022 · 196 citations
- Privacy-Preserving Face Recognition in the Frequency DomainYinggui Wang, Jian Liu, Man Luo, Le Yang et al.AAAI 2022 · 62 citations
Builds on6
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 1,822 citations
- Mixup Inference: Better Exploiting Mixup to Defend Adversarial AttacksTianyu Pang, Kun Xu, Jun ZhuICLR 2020 · 114 citations
- Learning mixtures of linear regressions in subexponential time via Fourier momentsSitan Chen, Jerry Li, Zhao SongSTOC 2020 · 16 citations
Related papers
- A Fusion-Denoising Attack on InstaHide with Data AugmentationXinjian Luo, Xiaokui Xiao, Yuncheng Wu, Juncheng Liu et al.AAAI 2022 · 9 citations
- Is Private Learning Possible with Instance Encoding?Nicholas Carlini, Samuel Deng, Sanjam Garg, Somesh Jha et al.S&P 2021 · 45 citations
- On InstaHide, Phase Retrieval, and Sparse Matrix FactorizationSitan Chen, Xiaoxiao Li, Zhao Song, Danyang ZhuoICLR 2021 · 1 citation
- When Deep Learning Meets Steganography: Protecting Inference Privacy in the DarkQin Liu, Jiamin Yang, Hongbo Jiang, Jie Wu et al.INFOCOM 2022 · 8 citations
- Deep Models Under the GAN: Information Leakage from Collaborative Deep LearningBriland Hitaj, Giuseppe Ateniese, Fernando Pérez-CruzCCS 2017 · 1,581 citations
