LAMP: Extracting Text from Gradients with Language Model Priors
Mislav Balunovic, Dimitar I. Dimitrov, Nikola Jovanovic, Martin T. Vechev
摘要
Recent work shows that sensitive user data can be reconstructed from gradient updates, breaking the key privacy promise of federated learning. While success was demonstrated primarily on image data, these methods do not directly transfer to other domains such as text. In this work, we propose LAMP, a novel attack tailored to textual data, that successfully reconstructs original text from gradients. Our attack is based on two key insights: (i) modeling prior text probability with an auxiliary language model, guiding the search towards more natural text, and (ii) alternating continuous and discrete optimization, which minimizes reconstruction loss on embeddings, while avoiding local minima by applying discrete text transformations. Our experiments demonstrate that LAMP is significantly more effective than prior work: it reconstructs 5x more bigrams and 23% longer subsequences on average. Moreover, we are the first to recover inputs from batch sizes larger than 1 for textual models. These findings indicate that gradient updates of models operating on textual data leak more information than previously thought.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed InferencesMengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li 等USENIX ATC 2024 · 被引用 78 次
- SPEAR: Exact Gradient Inversion of Batches in Federated LearningDimitar I. Dimitrov, Maximilian Baader, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 被引用 29 次
- DAGER: Exact Gradient Inversion for Large Language ModelsIvo Petrov, Dimitar I. Dimitrov, Maximilian Baader, Mark Niklas Müller 等NeurIPS 2024 · 被引用 29 次
- Federated Few-Shot Learning for Mobile NLPDongqi Cai, Shangguang Wang, Yaozong Wu, Felix Xiaozhu Lin 等MobiCom 2023 · 被引用 28 次
- Inf2Guard: An Information-Theoretic Framework for Learning Privacy-Preserving Representations against Inference AttacksSayedeh Leila Noorbakhsh, Binghui Zhang, Yuan Hong, Binghui WangUSENIX Security 2024 · 被引用 17 次
它引用的顶会 Paper10
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Inverting Gradients - How easy is it to break privacy in federated learning?Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, Michael MoellerNeurIPS 2020 · 被引用 1,822 次
- Evaluating Gradient Inversion Attacks and Defenses in Federated LearningYangsibo Huang, Samyak Gupta, Zhao Song, Kai Li 等NeurIPS 2021 · 被引用 419 次
- Gradient Inversion with Generative Image PriorJinwoo Jeon, Jaechang Kim, Kangwook Lee, Sewoong Oh 等NeurIPS 2021 · 被引用 216 次
- R-GAP: Recursive Gradient Attack on PrivacyJunyi Zhu, Matthew B. BlaschkoICLR 2021 · 被引用 157 次
相关 Paper
- Recovering Private Text in Federated Learning of Language ModelsSamyak Gupta, Yangsibo Huang, Zexuan Zhong, Tianyu Gao 等NeurIPS 2022 · 被引用 120 次
- Decepticons: Corrupted Transformers Breach Privacy in Federated Learning for Language ModelsLiam H. Fowl, Jonas Geiping, Steven Reich, Yuxin Wen 等ICLR 2023 · 被引用 10 次
- Panning for Gold in Federated Learning: Targeted Text Extraction under Arbitrarily Large-Scale AggregationHong-Min Chu, Jonas Geiping, Liam H. Fowl, Micah Goldblum 等ICLR 2023
- Reconstructing Training Data from Adapter-based Federated Large Language ModelsSilong Chen, Yuchuan Luo, Guilin Deng, Yi Liu 等WWW 2026
- Uncovering Gradient Inversion Risks in Practical Language Model TrainingXinguo Feng, Zhongkui Ma, Zihan Wang, Eu Joe Chegne 等CCS 2024 · 被引用 2 次
