From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models
Zixuan GU, Xiaojun Ye, Yang Liu
摘要
Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure through external APIs against the high computational cost of local deployment. Split learning has therefore emerged as a promising paradigm for LLM fine-tuning and inference under limited local resources. However, it introduces new privacy risks. Prior work primarily studies leakage of private input prompts, typically via inversion attacks on intermediate representations, while the potential for sensitive information leakage through generative response outputs remains largely unexplored. In this work, we unveil novel vulnerabilities of Split-LLM by presenting P atched Model I nversion with D ual-Sided I nitialization( PIDI ), a two-stage attack that simultaneously targets both private input prompts and output responses in Split-LLM settings. It combines dual-sided initialization with a patched inversion strategy to tackle long sequences, substantially outperforming prior inversion methods. To counter threats from both sides, we further propose the A dapter-based D ualGuard with M utual I nformation Defense( ADMI ), which integrates an adapter-based local warmup strategy and mutual information regularization to provide a strong empirical privacy protection with minimal impact on task performance. Extensive experiments across diverse tasks and models demonstrate that ADMI effectively defends against PIDI and other state-of-the-art inversion attacks. Our code is publicly available at https://github.com/FLAIR-THU/VFLAIR-LLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 被引用 200 次
- Label Leakage and Protection in Two-party Split LearningOscar Li, Jiankai Sun, Xin Yang, Weihao Gao 等ICLR 2022 · 被引用 170 次
- LAMP: Extracting Text from Gradients with Language Model PriorsMislav Balunovic, Dimitar I. Dimitrov, Nikola Jovanovic, Martin T. VechevNeurIPS 2022 · 被引用 100 次
- Split-and-Denoise: Protect large language model inference with local differential privacyPeihua Mai, Ran Yan, Zhe Huang, Youjia Yang 等ICML 2024 · 被引用 41 次
相关 Paper
- DualGuard: A Parameter Space Transformation Approach for Bidirectional Defense in Split-Based LLM Fine-TuningZihan Liu, Yizhen Wang, Rui Wang, Sai WuACL 2025
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
- Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced AttackGuanzhong Chen, Zhenghan Qin, Mingxin Yang, Yajie Zhou 等CCS 2024 · 被引用 7 次
- PISA: Privacy-Preserving Split Adaptation with Model IP ProtectionHaocheng Yang, Xiang Cheng, ZONGDA HAN, Pengjie Wang 等ICML 2026
- SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference AttacksKaiyuan Zhang, Siyuan Cheng, Hanxi Guo, Yuetian Chen 等USENIX Security 2025
