CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset Training
Qi Li, Cheng-Long Wang, Yinzhi Cao, Di Wang
摘要
Training models on a carefully chosen portion of data rather than the full dataset is now a standard preprocess for modern ML. From vision coreset selection to large-scale filtering in language models, it enables scalability with minimal utility loss. A common intuition is that training on fewer samples should also reduce privacy risks. In this paper, we challenge this assumption. We show that subset training is not privacy free: the very choices of which data are included or excluded can introduce new privacy surface and leak more sensitive information. Such information can be captured by adversaries either through side-channel metadata from the subset selection process or via the outputs of the target model. To systematically study this phenomenon, we propose CoLA (Choice Leakage Attack), a unified framework for analyzing privacy leakage in subset selection. In CoLA, depending on the adversary's knowledge of the side-channel information, we define two practical attack scenarios: Subset-aware Side-channel Attacks and Black-box Attacks. Under both scenarios, we investigate two privacy surfaces unique to subset training: (1) Training-membership MIA (TM-MIA), which concerns only the privacy of training data membership, and (2) Selection-participation MIA (SP-MIA), which concerns the privacy of all samples that participated in the subset selection process. Notably, SP-MIA enlarges the notion of membership from model training to the entire data-model supply chain. Experiments on vision and language models show that existing threat models underestimate subset-training privacy risks: the expanded privacy surface leaks both training and selection membership, extending risks from individual models to the broader ML ecosystem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang 等NDSS 2019 · 被引用 1,141 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
相关 Paper
- PANORAMIA: Privacy Auditing of Machine Learning Models without RetrainingMishaal Kazmi, Hadrien Lautraite, Alireza Akbari, Qiaoyue Tang 等NeurIPS 2024 · 被引用 26 次
- MIST: Defending Against Membership Inference Attacks Through Membership-Invariant Subspace TrainingJiacheng Li, Ninghui Li, Bruno RibeiroUSENIX Security 2024 · 被引用 16 次
- United We Defend: Collaborative Membership Inference Defenses in Federated LearningLi Bai, Junxu Liu, Sen Zhang, Xinwei Zhang 等USENIX Security 2026
- Free Record-Level Privacy Risk Evaluation Through Artifact-Based MethodsJoseph Pollock, Igor Shilov, Euodia Dodd, Yves-Alexandre de MontjoyeUSENIX Security 2025
- Privacy Side Channels in Machine Learning SystemsEdoardo Debenedetti, Giorgio Severi, Milad Nasr, Christopher A. Choquette-Choo 等USENIX Security 2024 · 被引用 52 次
