DualGuard: A Parameter Space Transformation Approach for Bidirectional Defense in Split-Based LLM Fine-Tuning
Zihan Liu, Yizhen Wang, Rui Wang, Sai Wu
摘要
Integrating split learning with large language model fine-tuning (LLM-FT) enables secure collaboration between a trusted local client and a well-equipped remote server, but it is vulnerable to data reconstruction attacks (DRAs) that exploit transmitted activations and gradients. Current defense methods, like adding noise to activations or gradients, often sacrifice task-specific model performance under strict privacy constraints. This paper introduces Du-alGuard, a bidirectional defense mechanism against DRAs for split-based LLM-FT. Du-alGuard proposes a local warm-up parameter space transformation to alter client-side model parameters before training, using multi-task learning to strike a balance between privacy protection and model performance. Additionally, a global fine-tuning parameter space retention strategy prevents the model from reverting to vulnerable states during formal fine-tuning. Experiments show that DualGuard outperforms current defense methods against various DRAs, while maintaining task performance. Our code will be made publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Split-and-Denoise: Protect large language model inference with local differential privacyPeihua Mai, Ran Yan, Zhe Huang, Youjia Yang 等ICML 2024 · 被引用 41 次
相关 Paper
- Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced AttackGuanzhong Chen, Zhenghan Qin, Mingxin Yang, Yajie Zhou 等CCS 2024 · 被引用 7 次
- From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language ModelsZixuan GU, Xiaojun Ye, Yang LiuICML 2026
- InfoDecom: Decomposing Information for Defending Against Privacy Leakage in Split InferenceRuijun Deng, Zhihui Lu, Qiang DuanAAAI 2026
- Focusing on Pinocchio's Nose: A Gradients Scrutinizer to Thwart Split-Learning Hijacking Attacks Using Intrinsic AttributesJiayun Fu, Xiaojing Ma, Bin B. Zhu, Pingyi Hu 等NDSS 2023
- PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud CollaborationYi Liu, Weixiang Han, Chengjun Cai, Xingliang Yuan 等INFOCOM 2026 · 被引用 2 次
