A Model-Agnostic Approach to Differentially Private Topic Mining
Han Wang, Jayashree Sharma, Shuya Feng, Kai Shu, Yuan Hong
摘要
Topic mining extracts patterns and insights from text data (e.g., documents, emails and product reviews), which can be used in various applications such as intent detection. However, topic mining can result in severe privacy threats to the users who have contributed to the text corpus since they can be re-identified from the text data with certain background knowledge. To our best knowledge, we propose the first differentially private topic mining technique (namely TopicDP) which injects well-calibrated Gaussian noise into the matrix output of any topic mining algorithm to ensure differential privacy and good utility. Specifically, we smoothen the sensitivity for the Gaussian mechanism via sensitivity sampling, which addresses the major challenges resulted from the high sensitivity in topic mining for differential privacy. Furthermore, we theoretically prove the differential privacy guarantee under the Rényi differential privacy mechanism and the utility error bounds of TopicDP. Finally, we conduct extensive experiments on two real-word text datasets (Enron email and Amazon Reviews), and the experimental results demonstrate that TopicDP is a model-agnostic framework that can generate better privacy preserving performance for topic mining as compared against other differential privacy mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Task-Agnostic Privacy-Preserving Representation Learning for Federated Learning against Attribute Inference AttacksCaridad Arroyo Arevalo, Sayedeh Leila Noorbakhsh, Yun Dong, Yuan Hong 等AAAI 2024 · 被引用 26 次
- Inf2Guard: An Information-Theoretic Framework for Learning Privacy-Preserving Representations against Inference AttacksSayedeh Leila Noorbakhsh, Binghui Zhang, Yuan Hong, Binghui WangUSENIX Security 2024 · 被引用 17 次
- DPI: Ensuring Strict Differential Privacy for Infinite Data StreamingShuya Feng, Meisam Mohammady, Han Wang, Xiaochen Li 等S&P 2024 · 被引用 17 次
它引用的顶会 Paper2
- MVG Mechanism: Differential Privacy under Matrix-Valued QueryThee Chanyaswad, Alex Dytso, H. Vincent Poor, Prateek MittalCCS 2018 · 被引用 55 次
- An end-to-end Differentially Private Latent Dirichlet Allocation Using a Spectral AlgorithmChris Decarolis, Mukul Ram, Seyed Esmaeili, Yu-Xiang Wang 等ICML 2020 · 被引用 12 次
相关 Paper
- Improving Sparse Vector Technique with Renyi Differential PrivacyYuqing Zhu, Yu-Xiang WangNeurIPS 2020 · 被引用 25 次
- Differentially Private n-gram ExtractionKunho Kim, Sivakanth Gopi, Janardhan Kulkarni, Sergey YekhaninNeurIPS 2021 · 被引用 22 次
- Sentence-level Privacy for Document EmbeddingsCasey Meehan, Khalil Mrini, Kamalika ChaudhuriACL 2022 · 被引用 26 次
- PrivateMail: Supervised Manifold Learning of Deep Features with Privacy for Image RetrievalPraneeth Vepakomma, Julia Balla, Ramesh RaskarAAAI 2022 · 被引用 4 次
- Federated Latent Dirichlet Allocation: A Local Differential Privacy Based FrameworkYansheng Wang, Yongxin Tong, Dingyuan ShiAAAI 2020 · 被引用 128 次
