Formal Privacy Proof of Data Encoding: The Possibility and Impossibility of Learnable Encryption
Hanshen Xiao, G. Edward Suh, Srinivas Devadas
Abstract
We initiate a formal study on the concept of learnable obfuscation and aim to answer the following question: is there a type of data encoding that maintains the "learnability" of encoded samples, thereby enabling direct model training on transformed data, while ensuring the privacy of both plaintext and the secret encoding function? This long-standing open problem has prompted many efforts to design such an encryption function, for example, Neu-raCrypt and TransNet. Nonetheless, all existing constructions are heuristic without formal privacy guarantees, and many successful reconstruction attacks are known on these constructions assuming an adversary with substantial prior knowledge. We present both generic possibility and impossibility results pertaining to learnable obfuscation. On one hand, we demonstrate that any non-trivial, property-preserving transformation which enables effectively learning over encoded samples cannot offer cryptographic computational security in the worst case. On the other hand, from the lens of information-theoretical security, we devise a series of new tools to produce provable and useful privacy guarantees from a set of heuristic obfuscation methods, including matrix masking, data mixing and permutation, through noise perturbation. Under the framework of PAC Privacy, we show how to quantify the leakage from the learnable obfuscation built upon obfuscation and perturbation methods against adversarial inference. Significantly sharpened utility-privacy tradeoffs are achieved compared to stateof-the-art accounting methods when measuring privacy against data reconstruction and membership inference attacks. CCS CONCEPTS • Security and privacy → Privacy-preserving protocols.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- BatchCrypt: Efficient Homomorphic Encryption for Cross-Silo Federated LearningChengliang Zhang, Suyi Li, Junzhe Xia, Wei Wang et al.USENIX ATC 2020 · 967 citations
- Differentially Private Learning Needs Better Features (or Much More Data)Florian Tramèr, Dan BonehICLR 2021 · 325 citations
Related papers
- When Does Data Augmentation Help With Membership Inference Attacks?Yigitcan Kaya, Tudor DumitrasICML 2021 · 82 citations
- Label-Only Membership Inference AttacksChristopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, Nicolas PapernotICML 2021 · 628 citations
- MemGuard: Defending against Black-Box Membership Inference Attacks via Adversarial ExamplesJinyuan Jia, Ahmed Salem, Michael Backes, Yang Zhang et al.CCS 2019 · 464 citations
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 543 citations
- RelaxLoss: Defending Membership Inference Attacks without Losing UtilityDingfan Chen, Ning Yu, Mario FritzICLR 2022 · 61 citations
