Split-and-Denoise: Protect large language model inference with local differential privacy
Peihua Mai, Ran Yan, Zhe Huang, Youjia Yang, Yan Pang
摘要
Large Language Models (LLMs) excel in natural language understanding by capturing hidden semantics in vector space. This process enriches the value of text embeddings for various downstream tasks, thereby fostering the Embedding-as-a-Service (EaaS) business model. However, the risk of privacy leakage due to direct text transmission to servers remains a critical concern. To address this, we introduce Split-N-Denoise (SnD), an private inference framework that splits the model to execute the token embedding layer on the client side at minimal computational cost. This allows the client to introduce noise prior to transmitting the embeddings to the server, and subsequently receive and denoise the perturbed output embeddings for downstream tasks. Our approach is designed for the inference stage of LLMs and requires no modifications to the model parameters. Extensive experiments demonstrate SnD's effectiveness in optimizing the privacy-utility tradeoff across various LLM architectures and diverse downstream tasks. The results reveal an improvement in performance under the same privacy budget compared to the baselines by over 10% on average, offering clients a privacy-preserving solution for local privacy protection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM InferenceJiangou Zhan, Wenhui Zhang, Zheng Zhang, Huanran Xue 等AAAI 2025 · 被引用 10 次
- A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language ModelsZuan Xie, Yang Xu, Hongli Xu, Yunming Liao 等INFOCOM 2026 · 被引用 10 次
- Personalized Federated Fine-Tuning for LLMs via Data-Driven Heterogeneous Model ArchitecturesYicheng Zhang, Zhen Qin, Zhaomin Wu, Jian Hou 等WWW 2026 · 被引用 9 次
- Prεεmpt: Sanitizing Sensitive Prompts for LLMsAmrita Roy Chowdhury, David Glukhov, Divyam Anshumaan, Prasad Chalasani 等NDSS 2026 · 被引用 5 次
- PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch CollaborationJunfei Zhan, Haoxun Shen, Zheng Lin, Tengjiao HeAAAI 2026 · 被引用 4 次
它引用的顶会 Paper8
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi 等ICLR 2022 · 被引用 494 次
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing 等NeurIPS 2022 · 被引用 209 次
- Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language ModelsHaonan Duan, Adam Dziedzic, Nicolas Papernot, Franziska BoenischNeurIPS 2023 · 被引用 116 次
- Bounding Training Data Reconstruction in Private (Deep) LearningChuan Guo, Brian Karrer, Kamalika Chaudhuri, Laurens van der MaatenICML 2022 · 被引用 66 次
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassMinxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang 等CCS 2023 · 被引用 35 次
相关 Paper
- HiddenEcho: Mitigating Noise Amplification in Differentially Private LLMs with Hidden-State CorrectionWenhao Li, Kunhao Li, Lei YangICLR 2026
- Reconstruction Attack-Resistant Inference Paradigm for LLM Cloud ServicesZipeng Ye, Wenjian Luo, Qi Zhou, Yubo TangAAAI 2026
- PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud CollaborationYi Liu, Weixiang Han, Chengjun Cai, Xingliang Yuan 等INFOCOM 2026 · 被引用 2 次
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
- DP-Fusion: Token-Level Differentially Private Inference for Large Language ModelsRushil Thareja, Preslav Nakov, Praneeth Vepakomma, Nils LukasICLR 2026 · 被引用 7 次
