DataStealing: Steal Data from Diffusion Models in Federated Learning with Multiple Trojans
Yuan Gan, Jiaxu Miao, Yi Yang
Abstract
Federated Learning (FL) is commonly used to collaboratively train models with privacy preservation. In this paper, we found out that the popular diffusion models have introduced a new vulnerability to FL, which brings serious privacy threats. Despite stringent data management measures, attackers can steal massive private data from local clients through multiple Trojans, which control generative behaviors with multiple triggers. We refer to the new task as DataStealing and demonstrate that the attacker can achieve the purpose based on our proposed Combinatorial Triggers (ComboTs) in a vanilla FL system. However, advanced distance-based FL defenses are still effective in filtering the malicious update according to the distances between each local update. Hence, we propose an Adaptive Scale Critical Parameters (AdaSCP) attack to circumvent the defenses and seamlessly incorporate malicious updates into the global model. Specifically, AdaSCP evaluates the importance of parameters with the gradients in dominant timesteps of the diffusion model. Subsequently, it adaptively seeks the optimal scale factor and magnifies critical parameter updates before uploading to the server. As a result, the malicious update becomes similar to the benign update, making it difficult for distance-based defenses to identify. Extensive experiments reveal the risk of leaking thousands of images in training diffusion models with FL. Moreover, these experiments demonstrate the effectiveness of AdaSCP in defeating advanced distance-based defenses. We hope this work will attract more attention from the FL community to the critical privacy security issues of Diffusion Models. Code: https://github.com/yuangan/DataStealing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e1a28ef-c1ed-46a0-a4fe-7d2e0aa9a5f5Cited by top-tier papers1
Ask how each one uses itBuilds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Comprehensive Privacy Analysis of Deep Learning: Passive and Active White-box Inference Attacks against Centralized and Federated LearningMilad Nasr, Reza Shokri, Amir HoumansadrS&P 2019 · 1,778 citations
Related papers
- Good Gradients Poison Your Model: Evading Defenses in Federated Learning via Boundary-adaptive PerturbationXiaojie Zhao, Jinqiao Shi, Yi Li, Junmin Huang et al.AAAI 2026
- STDLens: Model Hijacking-Resilient Federated Learning for Object DetectionKa-Ho Chow, Ling Liu, Wenqi Wei, Fatih Ilhan et al.CVPR 2023
- Poisoning with a Pill: Circumventing Detection in Federated LearningHanxi Guo, Hao Wang, Tao Song, Tianhang Zheng et al.AAAI 2026
- FedInv: Byzantine-Robust Federated Learning by Inversing Local Model UpdatesBo Zhao, Peng Sun, Tao Wang, Keyu JiangAAAI 2022 · 82 citations
- Enhancing Privacy Preservation in Federated Learning via Learning Rate PerturbationGuangnian Wan, Haitao Du, Xuejing Yuan, Jun Yang et al.ICCV 2023 · 2 citations
