Decepticons: Corrupted Transformers Breach Privacy in Federated Learning for Language Models
Liam H. Fowl, Jonas Geiping, Steven Reich, Yuxin Wen, Wojciech Czaja, Micah Goldblum, Tom Goldstein
Abstract
Privacy is a central tenet of Federated learning (FL), in which a central server trains models without centralizing user data. However, gradient updates used in FL can leak user information. While the most industrial uses of FL are for text applications (e.g. keystroke prediction), the majority of attacks on user privacy in FL have focused on simple image classifiers and threat models that assume honest execution of the FL protocol from the server. We propose a novel attack that reveals private user text by deploying malicious parameter vectors, and which succeeds even with mini-batches, multiple users, and long sequences. Unlike previous attacks on FL, the attack exploits characteristics of both the Transformer architecture and the token embedding, separately extracting tokens and positional embeddings to retrieve high-fidelity text. We argue that the threat model of malicious server states is highly relevant from a user-centric perspective, and show that in this scenario, text applications using transformer models are much more vulnerable than previously thought. * Authors contributed equally. Order chosen randomly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef8bff1a-cac6-4a4b-8bb1-1caff4d6e98aCited by top-tier papers19
- LAMP: Extracting Text from Gradients with Language Model PriorsMislav Balunovic, Dimitar I. Dimitrov, Nikola Jovanovic, Martin T. VechevNeurIPS 2022 · 100 citations
- Truth Serum: Poisoning Machine Learning Models to Reveal Their SecretsFlorian Tramèr, Reza Shokri, Ayrton San Joaquin, Hoang Le et al.CCS 2022 · 55 citations
- Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained ModelsYuxin Wen, Leo Marchyok, Sanghyun Hong, Jonas Geiping et al.NeurIPS 2024 · 39 citations
- SPEAR: Exact Gradient Inversion of Batches in Federated LearningDimitar I. Dimitrov, Maximilian Baader, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 29 citations
- DAGER: Exact Gradient Inversion for Large Language ModelsIvo Petrov, Dimitar I. Dimitrov, Maximilian Baader, Mark Niklas Müller et al.NeurIPS 2024 · 29 citations
Builds on14
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone et al.CCS 2017 · 3,936 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 1,822 citations
- Exploiting Unintended Feature Leakage in Collaborative LearningLuca Melis, Congzheng Song, Emiliano De Cristofaro, Vitaly ShmatikovS&P 2019 · 1,736 citations
- Evaluating Differentially Private Machine Learning in PracticeBargav Jayaraman, David EvansUSENIX Security 2019 · 586 citations
Related papers
- Panning for Gold in Federated Learning: Targeted Text Extraction under Arbitrarily Large-Scale AggregationHong-Min Chu, Jonas Geiping, Liam H. Fowl, Micah Goldblum et al.ICLR 2023
- Robbing the Fed: Directly Obtaining Private Data in Federated Learning with Modified ModelsLiam H. Fowl, Jonas Geiping, Wojciech Czaja, Micah Goldblum et al.ICLR 2022 · 181 citations
- Fishing for User Data in Large-Batch Federated Learning via Gradient MagnificationYuxin Wen, Jonas Geiping, Liam Fowl, Micah Goldblum et al.ICML 2022 · 119 citations
- Gradient Disaggregation: Breaking Privacy in Federated Learning by Reconstructing the User Participant MatrixMaximilian Lam, Gu-Yeon Wei, David Brooks, Vijay Janapa Reddi et al.ICML 2021 · 78 citations
- Gradient Inversion with Generative Image PriorJinwoo Jeon, Jaechang Kim, Kangwook Lee, Sewoong Oh et al.NeurIPS 2021 · 216 citations
