Sentence-level Privacy for Document Embeddings
Casey Meehan, Khalil Mrini, Kamalika Chaudhuri
Abstract
User language data can contain highly sensitive personal content. As such, it is imperative to offer users a strong and interpretable privacy guarantee when learning from their data. In this work we propose SentDP, pure local differential privacy at the sentence level for a single user document. We propose a novel technique, DeepCandidate, that combines concepts from robust statistics and language modeling to produce high (768) dimensional, general -SentDP document embeddings. This guarantees that any single sentence in a document can be substituted with any other sentence while keeping the embedding -indistinguishable. Our experiments indicate that these private document embeddings are useful for downstream tasks like sentiment analysis and topic classification and even outperform baseline methods with weaker guarantees like word-level Metric DP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassMinxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang et al.CCS 2023 · 35 citations
- Sanitizing Sentence Embeddings (and Labels) for Local Differential PrivacyMinxin Du, Xiang Yue, Sherman S. M. Chow, Huan SunWWW 2023 · 26 citations
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar et al.ACL 2023 · 24 citations
- Reconstructing training data from document understanding modelsJérémie Dentan, Arnaud Paran, Aymen ShabouUSENIX Security 2024 · 3 citations
Builds on4
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 291 citations
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
- Recursive Tree-Structured Self-Attention for Answer Sentence SelectionKhalil Mrini, Emilia Farcas, Ndapa NakasholeACL 2021
Related papers
- Private Language Models via Truncated Laplacian MechanismTianhao Huang, Tao Yang, Ivan Habernal, Lijie Hu et al.EMNLP 2024
- Learning to Generate Image Embeddings with User-Level Differential PrivacyZheng Xu, Maxwell D. Collins, Yuxiao Wang, Liviu Panait et al.CVPR 2023
- Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy GuaranteesStephen Meisenbacher, Maulik Chevli, Florian MatthesEMNLP 2025
- Concept-Aware Privacy Mechanisms for Defending Embedding Inversion AttacksYu-Che Tsai, Hsiang Hsiao, Kuan-Yu Chen, Shou-De LinICLR 2026 · 2 citations
- Structure-Preference Enabled Graph Embedding Generation Under Differential PrivacySen Zhang, Qingqing Ye, Haibo HuICDE 2025 · 1 citation
