Privacy-Preserving Data Deduplication for Enhancing Federated Learning of Language Models
Aydin Abadi, Vishnu Asutosh Dasu, Sumanta Sarkar
摘要
Deduplication is a vital preprocessing step that enhances machine learning model performance and saves training time and energy. However, enhancing federated learning through deduplication poses challenges, especially regarding scalability and potential privacy violations if deduplication involves sharing all clients' data. In this paper, we address the problem of deduplication in a federated setup by introducing a pioneering protocol, Efficient Privacy-Preserving Multi-Party Deduplication (EP-MPD). It efficiently removes duplicates from multiple clients' datasets without compromising data privacy. EP-MPD is constructed in a modular fashion, utilizing two novel variants of the Private Set Intersection protocol. Our extensive experiments demonstrate the significant benefits of deduplication in federated learning of large language models. For instance, we observe up to 19.62% improvement in perplexity and up to 27.95% reduction in running time while varying the duplication level between 10% and 30%. EP-MPD effectively balances privacy and performance in federated learning, making it a valuable solution for large-scale applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language ModelsPukang Ye, Junwei Luo, Jiachen Shen, Saipan Zhou 等NeurIPS 2025 · 被引用 1 次
- Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis]Jinming Hu, Jiahao Gu, Kenta Ploch, Hao Wang 等SIGMOD 2026
它引用的顶会 Paper13
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone 等CCS 2017 · 被引用 3,936 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 被引用 395 次
相关 Paper
- Privacy and Accuracy-Aware AI/ML Model DeduplicationHong Guan, Lei Yu, Lixi Zhou, Li Xiong 等SIGMOD 2025 · 被引用 3 次
- Data Duplication: A Novel Multi-Purpose Attack Paradigm in Machine UnlearningDayong Ye, Tianqing Zhu, Jiayang Li, Kun Gao 等USENIX Security 2025
- FedADMM: A Robust Federated Deep Learning Framework with Adaptivity to System HeterogeneityYonghai Gong, Yichuan Li, Nikolaos M. FrerisICDE 2022 · 被引用 41 次
- Efficient Differentially Private Secure Aggregation for Federated Learning via Hardness of Learning with ErrorsTimothy Stevens, Christian Skalka, Christelle Vincent, John H. Ring 等USENIX Security 2022
- SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-trainingNan He, Weichen Xiong, Hanwen Liu, Yi Liao 等ACL 2024 · 被引用 1 次
