Sentence-level Privacy for Document Embeddings
Casey Meehan, Khalil Mrini, Kamalika Chaudhuri
摘要
User language data can contain highly sensitive personal content. As such, it is imperative to offer users a strong and interpretable privacy guarantee when learning from their data. In this work we propose SentDP, pure local differential privacy at the sentence level for a single user document. We propose a novel technique, DeepCandidate, that combines concepts from robust statistics and language modeling to produce high (768) dimensional, general -SentDP document embeddings. This guarantees that any single sentence in a document can be substituted with any other sentence while keeping the embedding -indistinguishable. Our experiments indicate that these private document embeddings are useful for downstream tasks like sentiment analysis and topic classification and even outperform baseline methods with weaker guarantees like word-level Metric DP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassMinxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang 等CCS 2023 · 被引用 35 次
- Sanitizing Sentence Embeddings (and Labels) for Local Differential PrivacyMinxin Du, Xiang Yue, Sherman S. M. Chow, Huan SunWWW 2023 · 被引用 26 次
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar 等ACL 2023 · 被引用 24 次
- Reconstructing training data from document understanding modelsJérémie Dentan, Arnaud Paran, Aymen ShabouUSENIX Security 2024 · 被引用 3 次
它引用的顶会 Paper4
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 被引用 200 次
- Recursive Tree-Structured Self-Attention for Answer Sentence SelectionKhalil Mrini, Emilia Farcas, Ndapa NakasholeACL 2021
相关 Paper
- Private Language Models via Truncated Laplacian MechanismTianhao Huang, Tao Yang, Ivan Habernal, Lijie Hu 等EMNLP 2024
- Learning to Generate Image Embeddings with User-Level Differential PrivacyZheng Xu, Maxwell D. Collins, Yuxiao Wang, Liviu Panait 等CVPR 2023
- Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy GuaranteesStephen Meisenbacher, Maulik Chevli, Florian MatthesEMNLP 2025
- Concept-Aware Privacy Mechanisms for Defending Embedding Inversion AttacksYu-Che Tsai, Hsiang Hsiao, Kuan-Yu Chen, Shou-De LinICLR 2026 · 被引用 2 次
- Structure-Preference Enabled Graph Embedding Generation Under Differential PrivacySen Zhang, Qingqing Ye, Haibo HuICDE 2025 · 被引用 1 次
