Overlearning Reveals Sensitive Attributes
Congzheng Song, Vitaly Shmatikov
Abstract
"Overlearning" means that a model trained for a seemingly simple objective implicitly learns to recognize attributes and concepts that are (1) not part of the learning objective, and (2) sensitive from a privacy or bias perspective. For example, a binary gender classifier of facial images also learns to recognize raceseven races that are not represented in the training dataand identities. We demonstrate overlearning in several vision and NLP models and analyze its harmful consequences. First, inference-time representations of an overlearned model reveal sensitive attributes of the input, breaking privacy protections such as model partitioning. Second, an overlearned model can be "re-purposed" for a different, privacy-violating task even in the absence of the original training data. We show that overlearning is intrinsic for some tasks and cannot be prevented by censoring unwanted attributes. Finally, we investigate where, when, and why overlearning happens during model training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc9da73c-0810-4d3a-8600-24739722ddafCited by top-tier papers36
- The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure AggregationPeter Kairouz, Ziyu Liu, Thomas SteinkeICML 2021 · 291 citations
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
- Dataset Inference: Ownership Resolution in Machine LearningPratyush Maini, Mohammad Yaghini, Nicolas PapernotICLR 2021 · 155 citations
- When Machine Unlearning Jeopardizes PrivacyMin Chen, Zhikun Zhang, Tianhao Wang, Michael Backes et al.CCS 2021 · 146 citations
- Leakage of Dataset Properties in Multi-Party Machine LearningWanrong Zhang, Shruti Tople, Olga OhrimenkoUSENIX Security 2021 · 92 citations
Builds on1
Related papers
- Parameters or Privacy: A Provable Tradeoff Between Overparameterization and Membership InferenceJasper Tan, Blake Mason, Hamid Javadi, Richard G. BaraniukNeurIPS 2022 · 22 citations
- Overwriting Pretrained Bias with Finetuning DataAngelina Wang, Olga RussakovskyICCV 2023 · 50 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Making Users Indistinguishable: Attribute-wise Unlearning in Recommender SystemsYuyuan Li, Chaochao Chen, Xiaolin Zheng, Yizhao Zhang et al.ACM MM 2023 · 26 citations
- Does learning require memorization? a short tale about a long tailVitaly FeldmanSTOC 2020 · 28 citations
