Privacy-Preserving Data Deduplication for Enhancing Federated Learning of Language Models
Aydin Abadi, Vishnu Asutosh Dasu, Sumanta Sarkar
Abstract
Deduplication is a vital preprocessing step that enhances machine learning model performance and saves training time and energy. However, enhancing federated learning through deduplication poses challenges, especially regarding scalability and potential privacy violations if deduplication involves sharing all clients' data. In this paper, we address the problem of deduplication in a federated setup by introducing a pioneering protocol, Efficient Privacy-Preserving Multi-Party Deduplication (EP-MPD). It efficiently removes duplicates from multiple clients' datasets without compromising data privacy. EP-MPD is constructed in a modular fashion, utilizing two novel variants of the Private Set Intersection protocol. Our extensive experiments demonstrate the significant benefits of deduplication in federated learning of large language models. For instance, we observe up to 19.62% improvement in perplexity and up to 27.95% reduction in running time while varying the duplication level between 10% and 30%. EP-MPD effectively balances privacy and performance in federated learning, making it a valuable solution for large-scale applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 137d9be9-4f9b-49b2-86f2-37c9e05ee249Cited by top-tier papers2
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language ModelsPukang Ye, Junwei Luo, Jiachen Shen, Saipan Zhou et al.NeurIPS 2025 · 1 citation
- Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]Jinming Hu, Jiahao Gu, Kenta Ploch, Hao Wang et al.SIGMOD 2026
Builds on13
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone et al.CCS 2017 · 3,936 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
Related papers
- Privacy and Accuracy-Aware AI/ML Model DeduplicationHong Guan, Lei Yu, Lixi Zhou, Li Xiong et al.SIGMOD 2025 · 3 citations
- Data Duplication: A Novel Multi-Purpose Attack Paradigm in Machine UnlearningDayong Ye, Tianqing Zhu, Jiayang Li, Kun Gao et al.USENIX Security 2025
- FedADMM: A Robust Federated Deep Learning Framework with Adaptivity to System HeterogeneityYonghai Gong, Yichuan Li, Nikolaos M. FrerisICDE 2022 · 41 citations
- Efficient Differentially Private Secure Aggregation for Federated Learning via Hardness of Learning with ErrorsTimothy Stevens, Christian Skalka, Christelle Vincent, John H. Ring et al.USENIX Security 2022
- SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-trainingNan He, Weichen Xiong, Hanwen Liu, Yi Liao et al.ACL 2024 · 1 citation
