Federated Latent Dirichlet Allocation: A Local Differential Privacy Based Framework
Yansheng Wang, Yongxin Tong, Dingyuan Shi
Abstract
Latent Dirichlet Allocation (LDA) is a widely adopted topic model for industrial-grade text mining applications. However, its performance heavily relies on the collection of large amount of text data from users' everyday life for model training. Such data collection risks severe privacy leakage if the data collector is untrustworthy. To protect text data privacy while allowing accurate model training, we investigate federated learning of LDA models. That is, the model is collaboratively trained between an untrustworthy data collector and multiple users, where raw text data of each user are stored locally and not uploaded to the data collector. To this end, we propose FedLDA, a local differential privacy (LDP) based framework for federated learning of LDA models. Central in FedLDA is a novel LDP mechanism called Random Response with Priori (RRP), which provides theoretical guarantees on both data privacy and model accuracy. We also design techniques to reduce the communication cost between the data collector and the users during model training. Extensive experiments on three open datasets verified the effectiveness of our solution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfc930f0-8bef-43e9-985b-bfa76312032cCited by top-tier papers9
- Provably Secure Federated Learning against Malicious ClientsXiaoyu Cao, Jinyuan Jia, Neil Zhenqiang GongAAAI 2021 · 161 citations
- Hierarchical Personalized Federated Learning for User ModelingJinze Wu, Qi Liu, Zhenya Huang, Yuting Ning et al.WWW 2021 · 97 citations
- An Efficient Approach for Cross-Silo Federated Learning to RankYansheng Wang, Yongxin Tong, Dingyuan Shi, Ke XuICDE 2021 · 36 citations
- Distribution-Regularized Federated Learning on Non-IID DataYansheng Wang, Yongxin Tong, Zimu Zhou, Ruisheng Zhang et al.ICDE 2023 · 31 citations
- Prompt-enhanced Federated Content Representation Learning for Cross-domain RecommendationLei Guo, Ziang Lu, Junliang Yu, Quoc Viet Hung Nguyen et al.WWW 2024 · 30 citations
Builds on3
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone et al.CCS 2017 · 3,936 citations
- Heavy Hitter Estimation over Set-Valued Data with Local Differential PrivacyZhan Qin, Yin Yang, Ting Yu, Issa Khalil et al.CCS 2016 · 344 citations
- Locally Differentially Private Frequent Itemset MiningTianhao Wang, Ninghui Li, Somesh JhaS&P 2018 · 196 citations
Related papers
- LabelDP-Pro: Learning with Label Differential Privacy via ProjectionsBadih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar et al.ICLR 2024 · 4 citations
- Deep Learning with Label Differential PrivacyBadih Ghazi, Noah Golowich, Ravi Kumar, Pasin Manurangsi et al.NeurIPS 2021 · 193 citations
- Enhancing Privacy Preservation in Federated Learning via Learning Rate PerturbationGuangnian Wan, Haitao Du, Xuejing Yuan, Jun Yang et al.ICCV 2023 · 2 citations
- Efficient and Differentially Private Federated LLM Fine-Tuning on Heterogeneous ClientsNan Yan, Yuqing Li, Xiong Wang, Jing Chen et al.KDD 2026
- Differentially Private Federated Low Rank Adaptation Beyond Fixed-MatrixMing Wen, Jiaqi Zhu, Yuedong Xu, Yipeng Zhou et al.NeurIPS 2025 · 6 citations
