Lune

ICLR2025Top-tier venue

Privacy Auditing of Large Language Models

Ashwinee Panda, Xinyu Tang, Christopher A. Choquette-Choo, Milad Nasr, Prateek Mittal

2025Year
12Top-tier citations

Abstract

Current techniques for privacy auditing of large language models (LLMs) have limited efficacy-they rely on basic approaches to generate canaries which leads to weak membership inference attacks that in turn give loose lower bounds on the empirical privacy leakage. We develop canaries that are far more effective than those used in prior work under threat models that cover a range of realistic settings. We demonstrate through extensive experiments on multiple families of fine-tuned LLMs that our approach sets a new standard for detection of privacy leakage. For measuring the memorization rate of non-privately trained LLMs, our designed canaries surpass prior approaches. For example, on the Qwen2.5-0.5B model, our designed canaries achieve 49.6% TPR at 1% FPR, vastly surpassing the prior approach's 4.2% TPR at 1% FPR. Our method can be used to provide a privacy audit of ε ≈ 1 for a model trained with theoretical ε of 4. To the best of our knowledge, this is the first time that a privacy audit of LLM training has achieved nontrivial auditing success in the setting where the attacker cannot train shadow models, insert gradient canaries, or access the model at every iteration.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cac5e407-137b-4ebe-9c38-0a9982949cca

Cited by top-tier papers12

Ask how each one uses it

Builds on38

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines