Split-and-Denoise: Protect large language model inference with local differential privacy
Peihua Mai, Ran Yan, Zhe Huang, Youjia Yang, Yan Pang
Abstract
Large Language Models (LLMs) excel in natural language understanding by capturing hidden semantics in vector space. This process enriches the value of text embeddings for various downstream tasks, thereby fostering the Embedding-as-a-Service (EaaS) business model. However, the risk of privacy leakage due to direct text transmission to servers remains a critical concern. To address this, we introduce Split-N-Denoise (SnD), an private inference framework that splits the model to execute the token embedding layer on the client side at minimal computational cost. This allows the client to introduce noise prior to transmitting the embeddings to the server, and subsequently receive and denoise the perturbed output embeddings for downstream tasks. Our approach is designed for the inference stage of LLMs and requires no modifications to the model parameters. Extensive experiments demonstrate SnD's effectiveness in optimizing the privacy-utility tradeoff across various LLM architectures and diverse downstream tasks. The results reveal an improvement in performance under the same privacy budget compared to the baselines by over 10% on average, offering clients a privacy-preserving solution for local privacy protection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79cad940-e9cb-441d-8da4-bac6870396b4Cited by top-tier papers15
- Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM InferenceJiangou Zhan, Wenhui Zhang, Zheng Zhang, Huanran Xue et al.AAAI 2025 · 10 citations
- A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language ModelsZuan Xie, Yang Xu, Hongli Xu, Yunming Liao et al.INFOCOM 2026 · 10 citations
- Personalized Federated Fine-Tuning for LLMs via Data-Driven Heterogeneous Model ArchitecturesYicheng Zhang, Zhen Qin, Zhaomin Wu, Jian Hou et al.WWW 2026 · 9 citations
- Prεεmpt: Sanitizing Sensitive Prompts for LLMsAmrita Roy Chowdhury, David Glukhov, Divyam Anshumaan, Prasad Chalasani et al.NDSS 2026 · 5 citations
- PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch CollaborationJunfei Zhan, Haoxun Shen, Zheng Lin, Tengjiao HeAAAI 2026 · 4 citations
Builds on8
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- Iron: Private Inference on TransformersMeng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing et al.NeurIPS 2022 · 209 citations
- Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language ModelsHaonan Duan, Adam Dziedzic, Nicolas Papernot, Franziska BoenischNeurIPS 2023 · 116 citations
- Bounding Training Data Reconstruction in Private (Deep) LearningChuan Guo, Brian Karrer, Kamalika Chaudhuri, Laurens van der MaatenICML 2022 · 66 citations
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassMinxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang et al.CCS 2023 · 35 citations
Related papers
- HiddenEcho: Mitigating Noise Amplification in Differentially Private LLMs with Hidden-State CorrectionWenhao Li, Kunhao Li, Lei YangICLR 2026
- Reconstruction Attack-Resistant Inference Paradigm for LLM Cloud ServicesZipeng Ye, Wenjian Luo, Qi Zhou, Yubo TangAAAI 2026
- PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud CollaborationYi Liu, Weixiang Han, Chengjun Cai, Xingliang Yuan et al.INFOCOM 2026 · 2 citations
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
- DP-Fusion: Token-Level Differentially Private Inference for Large Language ModelsRushil Thareja, Preslav Nakov, Praneeth Vepakomma, Nils LukasICLR 2026 · 7 citations
