Optimizing Canaries for Privacy Auditing with Metagradient Descent
Matteo Boglioni, Terrance Liu, Andrew Ilyas, Steven Wu
Abstract
In this work we study black-box privacy auditing, where the goal is to lower bound the privacy parameter of a differentially private learning algorithm using only the algorithm’s outputs (i.e., final trained model). For DP-SGD (the most successful method for training differentially private deep learning models), the canonical auditing approach uses membership inference—an auditor comes with a small set of special “canary” examples, inserts a random subset of them into the training set, and then tries to discern which of their canaries were included in the training set (typically via a membership inference attack). The auditor’s success rate then provides a lower bound on the privacy parameters of the learning algorithm. Our main contribution is a method for optimizing the auditor’s canary set to improve privacy auditing, leveraging recent work on metagradient optimization (Engstrom et al., 2025). Our empirical evaluation demonstrates that in certain instances, using such optimized canaries can improve empirical lower bounds for differentially private image classification models by several times when compared to canaries proposed in prior work. Furthermore, we demonstrate that our method is DP-SGD agnostic and efficient: canaries optimized for non-private SGD with a small model architecture remain effective when auditing larger models trained with DP-SGD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aea20343-ddc4-4a53-acf9-7ce132424f8cCited by top-tier papers2
- OptiFluence: Principled Design of Privacy CanariesMohammad Yaghini, Michael Aerni, Junrui Zhang, Nicolas Papernot et al.ICML 2026
- Beyond Membership: Limitations of Add/Remove Adjacency in Differential PrivacyGauri Pradhan, Joonas Jälkö, Santiago Zanella-Béguelin, Antti HonkelaICLR 2026
Builds on19
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Auditing Differentially Private Machine Learning: How Private is Private SGD?Matthew Jagielski, Jonathan R. Ullman, Alina OpreaNeurIPS 2020 · 354 citations
- Adversary Instantiation: Lower Bounds for Differentially Private Machine LearningMilad Nasr, Shuang Song, Abhradeep Thakurta, Nicolas Papernot et al.S&P 2021 · 288 citations
- Stabilizing Differentiable Architecture Search via Perturbation-based RegularizationXiangning Chen, Cho-Jui HsiehICML 2020 · 235 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
Related papers
- Nearly Tight Black-Box Auditing of Differentially Private Machine LearningMeenatchi Sundaram Muthu Selva Annamalai, Emiliano De CristofaroNeurIPS 2024 · 32 citations
- Tight Auditing of Differentially Private Machine LearningMilad Nasr, Jamie Hayes, Thomas Steinke, Borja Balle et al.USENIX Security 2023
- Privacy Auditing of Large Language ModelsAshwinee Panda, Xinyu Tang, Christopher A. Choquette-Choo, Milad Nasr et al.ICLR 2025
- Bounding training data reconstruction in DP-SGDJamie Hayes, Borja Balle, Saeed MahloujifarNeurIPS 2023 · 73 citations
- Efficient Privacy Auditing for Generative Model via Local InformationJingnan Xu, Leixia Wang, Xiaofeng MengKDD 2026
