A Model-Agnostic Approach to Differentially Private Topic Mining
Han Wang, Jayashree Sharma, Shuya Feng, Kai Shu, Yuan Hong
Abstract
Topic mining extracts patterns and insights from text data (e.g., documents, emails and product reviews), which can be used in various applications such as intent detection. However, topic mining can result in severe privacy threats to the users who have contributed to the text corpus since they can be re-identified from the text data with certain background knowledge. To our best knowledge, we propose the first differentially private topic mining technique (namely TopicDP) which injects well-calibrated Gaussian noise into the matrix output of any topic mining algorithm to ensure differential privacy and good utility. Specifically, we smoothen the sensitivity for the Gaussian mechanism via sensitivity sampling, which addresses the major challenges resulted from the high sensitivity in topic mining for differential privacy. Furthermore, we theoretically prove the differential privacy guarantee under the Rényi differential privacy mechanism and the utility error bounds of TopicDP. Finally, we conduct extensive experiments on two real-word text datasets (Enron email and Amazon Reviews), and the experimental results demonstrate that TopicDP is a model-agnostic framework that can generate better privacy preserving performance for topic mining as compared against other differential privacy mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffe48f23-c199-4e39-9de9-08d16fc0b144Cited by top-tier papers3
- Task-Agnostic Privacy-Preserving Representation Learning for Federated Learning against Attribute Inference AttacksCaridad Arroyo Arevalo, Sayedeh Leila Noorbakhsh, Yun Dong, Yuan Hong et al.AAAI 2024 · 26 citations
- Inf2Guard: An Information-Theoretic Framework for Learning Privacy-Preserving Representations against Inference AttacksSayedeh Leila Noorbakhsh, Binghui Zhang, Yuan Hong, Binghui WangUSENIX Security 2024 · 17 citations
- DPI: Ensuring Strict Differential Privacy for Infinite Data StreamingShuya Feng, Meisam Mohammady, Han Wang, Xiaochen Li et al.S&P 2024 · 17 citations
Builds on2
- MVG Mechanism: Differential Privacy under Matrix-Valued QueryThee Chanyaswad, Alex Dytso, H. Vincent Poor, Prateek MittalCCS 2018 · 55 citations
- An end-to-end Differentially Private Latent Dirichlet Allocation Using a Spectral AlgorithmChris Decarolis, Mukul Ram, Seyed Esmaeili, Yu-Xiang Wang et al.ICML 2020 · 12 citations
Related papers
- Improving Sparse Vector Technique with Renyi Differential PrivacyYuqing Zhu, Yu-Xiang WangNeurIPS 2020 · 25 citations
- Differentially Private n-gram ExtractionKunho Kim, Sivakanth Gopi, Janardhan Kulkarni, Sergey YekhaninNeurIPS 2021 · 22 citations
- Sentence-level Privacy for Document EmbeddingsCasey Meehan, Khalil Mrini, Kamalika ChaudhuriACL 2022 · 26 citations
- PrivateMail: Supervised Manifold Learning of Deep Features with Privacy for Image RetrievalPraneeth Vepakomma, Julia Balla, Ramesh RaskarAAAI 2022 · 4 citations
- Federated Latent Dirichlet Allocation: A Local Differential Privacy Based FrameworkYansheng Wang, Yongxin Tong, Dingyuan ShiAAAI 2020 · 128 citations
