From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models
Zixuan GU, Xiaojun Ye, Yang Liu
Abstract
Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure through external APIs against the high computational cost of local deployment. Split learning has therefore emerged as a promising paradigm for LLM fine-tuning and inference under limited local resources. However, it introduces new privacy risks. Prior work primarily studies leakage of private input prompts, typically via inversion attacks on intermediate representations, while the potential for sensitive information leakage through generative response outputs remains largely unexplored. In this work, we unveil novel vulnerabilities of Split-LLM by presenting P atched Model I nversion with D ual-Sided I nitialization( PIDI ), a two-stage attack that simultaneously targets both private input prompts and output responses in Split-LLM settings. It combines dual-sided initialization with a patched inversion strategy to tackle long sequences, substantially outperforming prior inversion methods. To counter threats from both sides, we further propose the A dapter-based D ualGuard with M utual I nformation Defense( ADMI ), which integrates an adapter-based local warmup strategy and mutual information regularization to provide a strong empirical privacy protection with minimal impact on task performance. Extensive experiments across diverse tasks and models demonstrate that ADMI effectively defends against PIDI and other state-of-the-art inversion attacks. Our code is publicly available at https://github.com/FLAIR-THU/VFLAIR-LLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7f27164-8516-4481-bb70-6cda4c0fd2efBuilds on9
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 291 citations
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
- Label Leakage and Protection in Two-party Split LearningOscar Li, Jiankai Sun, Xin Yang, Weihao Gao et al.ICLR 2022 · 170 citations
- LAMP: Extracting Text from Gradients with Language Model PriorsMislav Balunovic, Dimitar I. Dimitrov, Nikola Jovanovic, Martin T. VechevNeurIPS 2022 · 100 citations
- Split-and-Denoise: Protect large language model inference with local differential privacyPeihua Mai, Ran Yan, Zhe Huang, Youjia Yang et al.ICML 2024 · 41 citations
Related papers
- DualGuard: A Parameter Space Transformation Approach for Bidirectional Defense in Split-Based LLM Fine-TuningZihan Liu, Yizhen Wang, Rui Wang, Sai WuACL 2025
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
- Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced AttackGuanzhong Chen, Zhenghan Qin, Mingxin Yang, Yajie Zhou et al.CCS 2024 · 7 citations
- PISA: Privacy-Preserving Split Adaptation with Model IP ProtectionHaocheng Yang, Xiang Cheng, ZONGDA HAN, Pengjie Wang et al.ICML 2026
- SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference AttacksKaiyuan Zhang, Siyuan Cheng, Hanxi Guo, Yuetian Chen et al.USENIX Security 2025
