PrE-Text: Training Language Models on Private Federated Data in the Age of LLMs
Charlie Hou, Akshat Shrivastava, Hongyuan Zhan, Rylan Conway, Trang Le, Adithya Sagar, Giulia Fanti, Daniel Lazar
摘要
On-device training is currently the most common approach for training machine learning (ML) models on private, distributed user data. Despite this, on-device training has several drawbacks: (1) most user devices are too small to train large models on-device, (2) on-device training is communication- and computation-intensive, and (3) on-device training can be difficult to debug and deploy. To address these problems, we propose Private Evolution-Text (PrE-Text), a method for generating differentially private (DP) synthetic textual data. First, we show that across multiple datasets, training small models (models that fit on user devices) with PrE-Text synthetic data outperforms small models trained on-device under practical privacy regimes (, ). We achieve these results while using 9 fewer rounds, 6 less client computation per round, and 100 less communication per round. Second, finetuning large models on PrE-Text's DP synthetic data improves large language model (LLM) performance on private data across the same range of privacy budgets. Altogether, these results suggest that training on DP synthetic data can be a better option than training a model on-device on private distributed data. Code is available at https://github.com/houcharlie/PrE-Text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM InferenceJiangou Zhan, Wenhui Zhang, Zheng Zhang, Huanran Xue 等AAAI 2025 · 被引用 10 次
- PrivCode: When Code Generation Meets Differential PrivacyZheng Liu, Chen Gong, Terry Yue Zhuo, Kecen Li 等NDSS 2026 · 被引用 5 次
- Private Evolution ConvergesTomás González Lara, Giulia Fanti, Aaditya RamdasNeurIPS 2025 · 被引用 3 次
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted ControlYuzheng Hu, Ryan McKenna, Da Yu, Shanshan Wu 等ICML 2026 · 被引用 1 次
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language ModelsPukang Ye, Junwei Luo, Jiachen Shen, Saipan Zhou 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper42
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi 等NeurIPS 2020 · 被引用 2,231 次
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 被引用 1,615 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- Differentially Private Synthetic Data via Foundation Model APIs 2: TextChulin Xie, Zinan Lin, Arturs Backurs, Sivakanth Gopi 等ICML 2024 · 被引用 71 次
- Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMsBowen Tan, Zheng Xu, Eric P. Xing, Zhiting Hu 等ICML 2025
- Synthetic Text Generation with Differential Privacy: A Simple and Practical RecipeXiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar 等ACL 2023 · 被引用 24 次
- Secret-Protected Evolution for Differentially Private Synthetic Text GenerationTianze Wang, Zhaoyu Chen, Jian Du, Yingtai Xiao 等ICLR 2026 · 被引用 1 次
- EPSVec: Efficient and Private Synthetic Data Generation via Dataset VectorsMohammadamin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong 等ICML 2026 · 被引用 1 次
