Scaling Laws for Differentially Private Language Models
Ryan McKenna, Yangsibo Huang, Amer Sinha, Borja Balle, Zachary Charles, Christopher A. Choquette-Choo, Badih Ghazi, Georgios Kaissis, Ravi Kumar, Ruibo Liu, Da Yu, Chiyuan Zhang
摘要
Scaling laws have emerged as important components of large language model (LLM) training as they can predict performance gains through scale, and provide guidance on important hyperparameter choices that would otherwise be expensive. LLMs also rely on large, high-quality training datasets, like those sourced from (sometimes sensitive) user data. Training models on this sensitive user data requires careful privacy protections like differential privacy (DP). However, the dynamics of DP training are significantly different, and consequently their scaling laws are not yet fully understood. In this work, we establish scaling laws that accurately model the intricacies of DP LLM training, providing a complete picture of the compute-privacy-utility tradeoffs and the optimal training configurations in many settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Efficient privacy loss accounting for subsampling and random allocationVitaly Feldman, Moshe ShenfeldICML 2026 · 被引用 5 次
- Fundamental Limitations of Favorable Privacy–Utility Guarantees for DP-SGDMurat Bilgehan Ertan), Marten van Dijk)CCS 2026 · 被引用 3 次
- Privasis: Synthesizing the Largest "Public" Private Dataset from ScratchHyunwoo Kim, Niloofar Mireshghallah, Michael Duan, Rui Xin 等ICML 2026 · 被引用 3 次
- On Optimal Hyperparameters for Differentially Private Deep Transfer LearningAki Rehn, Linzh Zhao, Mikko A. Heikkilä, Antti HonkelaICLR 2026 · 被引用 2 次
- Adaptive Methods Are Preferable in High Privacy Settings: An SDE PerspectiveEnea Monzio Compagnoni, Alessandro Stanghellini, Rustem Islamov, Aurélien Lucchi 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper29
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 被引用 502 次
相关 Paper
- DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language ModelsYanming Liu, Xinyue Peng, Yuwei Zhang, Xiaolan Ke 等AAAI 2025 · 被引用 4 次
- TAN Without a Burn: Scaling Laws of DP-SGDTom Sander, Pierre Stock, Alexandre SablayrollesICML 2023 · 被引用 58 次
- Differentially Private Model CompressionFatemehsadat Mireshghallah, Arturs Backurs, Huseyin A. Inan, Lukas Wutschitz 等NeurIPS 2022 · 被引用 18 次
- Open LLMs are Necessary for Current Private Adaptations and Outperform their Closed AlternativesVincent Hanke, Tom Blanchard, Franziska Boenisch, Iyiola E. Olatunji 等NeurIPS 2024 · 被引用 27 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
