Privacy Auditing of Large Language Models
Ashwinee Panda, Xinyu Tang, Christopher A. Choquette-Choo, Milad Nasr, Prateek Mittal
摘要
Current techniques for privacy auditing of large language models (LLMs) have limited efficacy-they rely on basic approaches to generate canaries which leads to weak membership inference attacks that in turn give loose lower bounds on the empirical privacy leakage. We develop canaries that are far more effective than those used in prior work under threat models that cover a range of realistic settings. We demonstrate through extensive experiments on multiple families of fine-tuned LLMs that our approach sets a new standard for detection of privacy leakage. For measuring the memorization rate of non-privately trained LLMs, our designed canaries surpass prior approaches. For example, on the Qwen2.5-0.5B model, our designed canaries achieve 49.6% TPR at 1% FPR, vastly surpassing the prior approach's 4.2% TPR at 1% FPR. Our method can be used to provide a privacy audit of ε ≈ 1 for a model trained with theoretical ε of 4. To the best of our knowledge, this is the first time that a privacy audit of LLM training has achieved nontrivial auditing success in the setting where the attacker cannot train shadow models, insert gradient canaries, or access the model at every iteration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Exploring the limits of strong membership inference attacks on large language modelsJamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski 等NeurIPS 2025 · 被引用 26 次
- Hubble: a Model Suite to Advance the Study of LLM MemorizationJohnny Wei, Ameya Godbole, Mohammad Aflah Khan, Ryan Yixiang Wang 等ICLR 2026 · 被引用 22 次
- Benchmarking Empirical Privacy Protection for Adaptations of Large Language ModelsBartlomiej Marek, Lorenzo Rossi, Vincent Hanke, Xun Wang 等ICLR 2026 · 被引用 8 次
- Optimizing Canaries for Privacy Auditing with Metagradient DescentMatteo Boglioni, Terrance Liu, Andrew Ilyas, Steven WuICLR 2026 · 被引用 7 次
- Train Once, Answer All: Many Pretraining Experiments for the Cost of OneSebastian Bordt, Martin PawelczykICLR 2026 · 被引用 6 次
它引用的顶会 Paper38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic TextMatthieu Meeus, Lukas Wutschitz, Santiago Zanella-Béguelin, Shruti Tople 等ICML 2025
- OptiFluence: Principled Design of Privacy CanariesMohammad Yaghini, Michael Aerni, Junrui Zhang, Nicolas Papernot 等ICML 2026
- Information-Theoretic Membership Inference for Granular Quantification of MemorizationJiashu Tao, Reza ShokriICLR 2026
- Powerful Training-Free Membership Inference Against Fine-Tuned Autoregressive Language ModelsDavid Ilic, David Stanojevic, Kostadin CvejoskiACL 2026
- Natural Identifiers for Privacy and Data Audits in Large Language ModelsLorenzo Rossi, Bartlomiej Marek, Franziska Boenisch, Adam DziedzicICLR 2026 · 被引用 3 次
