Overlearning Reveals Sensitive Attributes
Congzheng Song, Vitaly Shmatikov
摘要
"Overlearning" means that a model trained for a seemingly simple objective implicitly learns to recognize attributes and concepts that are (1) not part of the learning objective, and (2) sensitive from a privacy or bias perspective. For example, a binary gender classifier of facial images also learns to recognize raceseven races that are not represented in the training dataand identities. We demonstrate overlearning in several vision and NLP models and analyze its harmful consequences. First, inference-time representations of an overlearned model reveal sensitive attributes of the input, breaking privacy protections such as model partitioning. Second, an overlearned model can be "re-purposed" for a different, privacy-violating task even in the absence of the original training data. We show that overlearning is intrinsic for some tasks and cannot be prevented by censoring unwanted attributes. Finally, we investigate where, when, and why overlearning happens during model training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- The Distributed Discrete Gaussian Mechanism for Federated Learning with Secure AggregationPeter Kairouz, Ziyu Liu, Thomas SteinkeICML 2021 · 被引用 291 次
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 被引用 200 次
- Dataset Inference: Ownership Resolution in Machine LearningPratyush Maini, Mohammad Yaghini, Nicolas PapernotICLR 2021 · 被引用 155 次
- When Machine Unlearning Jeopardizes PrivacyMin Chen, Zhikun Zhang, Tianhao Wang, Michael Backes 等CCS 2021 · 被引用 146 次
- Leakage of Dataset Properties in Multi-Party Machine LearningWanrong Zhang, Shruti Tople, Olga OhrimenkoUSENIX Security 2021 · 被引用 92 次
它引用的顶会 Paper1
相关 Paper
- Parameters or Privacy: A Provable Tradeoff Between Overparameterization and Membership InferenceJasper Tan, Blake Mason, Hamid Javadi, Richard G. BaraniukNeurIPS 2022 · 被引用 22 次
- Overwriting Pretrained Bias with Finetuning DataAngelina Wang, Olga RussakovskyICCV 2023 · 被引用 50 次
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang 等ICCV 2019 · 被引用 469 次
- Making Users Indistinguishable: Attribute-wise Unlearning in Recommender SystemsYuyuan Li, Chaochao Chen, Xiaolin Zheng, Yizhao Zhang 等ACM MM 2023 · 被引用 26 次
- Does learning require memorization? a short tale about a long tailVitaly FeldmanSTOC 2020 · 被引用 28 次
