Uncovering Gradient Inversion Risks in Practical Language Model Training
Xinguo Feng, Zhongkui Ma, Zihan Wang, Eu Joe Chegne, Mengyao Ma, Alsharif Abuadbba, Guangdong Bai
Abstract
The gradient inversion attack has been demonstrated as a significant privacy threat to federated learning (FL), particularly in continuous domains such as vision models. In contrast, it is often considered less effective or highly dependent on impractical training settings when applied to language models, due to the challenges posed by the discrete nature of tokens in text data. As a result, its potential privacy threats remain largely underestimated, despite FL being an emerging training method for language models. In this work, we propose a domain-specific gradient inversion attack named Grab (gradient inversion with hybrid optimization). Grab features two alternating optimization processes to address the challenges caused by practical training settings, including a simultaneous optimization on dropout masks between layers for improved token recovery and a discrete optimization for effective token sequencing. Grab can recover a significant portion (up to 92.9% recovery rate) of the private training data, outperforming the attack strategy of utilizing discrete optimization with an auxiliary model by notable improvements of up to 28.9% recovery rate in benchmark settings and 48.5% recovery rate in practical settings. Grab provides a valuable step forward in understanding this privacy threat in the emerging FL training mode of language models. CCS Concepts • Security and privacy; • Computing methodologies → Machine learning;
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4465cc69-17fc-4da9-8cc7-ff80e050b8d5Cited by top-tier papers7
- SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value DecompositionChenxiang Luo, David K. Y. Yau, Qun SongNDSS 2026 · 3 citations
- Convex Hull Approximation for Activation FunctionsZhongkui Ma, Zihan Wang, Guangdong BaiOOPSLA 2025 · 2 citations
- ReTrace: Reinforcement Learning-Guided Reconstruction Attacks on Machine UnlearningMengyao Ma, Shuofeng Liu, Minhui Xue, Surya Nepal et al.ICLR 2026
- When the Aggregator Cheats: Data-Free Backdoors in Federated LLM-based QA SystemsChenqing Zhu, Yanbo Dai, Yulong Tian, Qingming Li et al.USENIX Security 2026
- Depth Gives a False Sense of Privacy: LLM Internal States InversionTian Dong, Yan Meng, Shaofeng Li, Guoxing Chen et al.USENIX Security 2025
Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- LAMP: Extracting Text from Gradients with Language Model PriorsMislav Balunovic, Dimitar I. Dimitrov, Nikola Jovanovic, Martin T. VechevNeurIPS 2022 · 100 citations
- Recovering Private Text in Federated Learning of Language ModelsSamyak Gupta, Yangsibo Huang, Zexuan Zhong, Tianyu Gao et al.NeurIPS 2022 · 120 citations
- Geminio: Language-Guided Gradient Inversion Attacks in Federated LearningJunjie Shan, Ziqi Zhao, Jialin Lu, Rui Zhang et al.ICCV 2025 · 2 citations
- DAGER: Exact Gradient Inversion for Large Language ModelsIvo Petrov, Dimitar I. Dimitrov, Maximilian Baader, Mark Niklas Müller et al.NeurIPS 2024 · 29 citations
- SoK: Gradient Inversion Attacks in Federated LearningVincenzo Carletti, Pasquale Foggia, Carlo Mazzocca, Giuseppe Parrella et al.USENIX Security 2025
