Use Privacy in Data-Driven Systems: Theory and Experiments with Machine Learnt Programs
Anupam Datta, Matthew Fredrikson, Gihyuk Ko, Piotr Mardziel, Shayak Sen
摘要
This paper presents an approach to formalizing and enforcing a class of use privacy properties in data-driven systems. In contrast to prior work, we focus on use restrictions on proxies (i.e. strong predictors) of protected information types. Our definition relates proxy use to intermediate computations that occur in a program, and identify two essential properties that characterize this behavior: 1) its result is strongly associated with the protected information type in question, and 2) it is likely to causally affect the final output of the program. For a specific instantiation of this definition, we present a program analysis technique that detects instances of proxy use in a model, and provides a witness that identifies which parts of the corresponding program exhibit the behavior. Recognizing that not all instances of proxy use of a protected information type are inappropriate, we make use of a normative judgment oracle that makes this inappropriateness determination for a given witness. Our repair algorithm uses the witness of an inappropriate proxy use to transform the model into one that provably does not exhibit proxy use, while avoiding changes that unduly affect classification accuracy. Using a corpus of social datasets, our evaluation shows that these algorithms are able to detect proxy use instances that would be difficult to find using existing techniques, and subsequently remove them while maintaining acceptable classification performance. CCS CONCEPTS • Security and privacy → Privacy protections; KEYWORDS use privacy INTRODUCTION Restrictions on information use occupy a central place in privacy regulations and legal frameworks [28, 54, 61, 62] . We introduce the term use privacy to refer to privacy norms governing information use. A number of recent cases have evidenced that inappropriate
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Quantitative Verification of Neural Networks and Its Security ApplicationsTeodora Baluta, Shiqi Shen, Shweta Shinde, Kuldeep S. Meel 等CCS 2019 · 被引用 115 次
- Perfectly parallel fairness certification of neural networksCaterina Urban, Maria Christakis, Valentin Wüstholz, Fuyuan ZhangOOPSLA 2020 · 被引用 61 次
- An Information-Theoretic Quantification of Discrimination with Exempt FeaturesSanghamitra Dutta, Praveen Venkatesh, Piotr Mardziel, Anupam Datta 等AAAI 2020 · 被引用 34 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- Fair Wrapping for Black-box PredictionsAlexander Soen, Ibrahim M. Alabdulmohsin, Sanmi Koyejo, Yishay Mansour 等NeurIPS 2022 · 被引用 8 次
它引用的顶会 Paper1
相关 Paper
- Weak Proxies are Sufficient and Preferable for Fairness with Missing Sensitive AttributesZhaowei Zhu, Yuanshun Yao, Jiankai Sun, Hang Li 等ICML 2023 · 被引用 28 次
- MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic PerspectiveYizhuo Chen, Chun-Fu Chen, Hsiang Hsu, Shaohan Hu 等ICML 2024 · 被引用 3 次
- Informational Friction as a Lens for Studying Algorithmic Aspects of PrivacyPatrick Skeba, Eric P. S. BaumerCSCW 2020 · 被引用 18 次
- The Feasibility of Dynamically Granted Permissions: Aligning Mobile Privacy with User PreferencesPrimal Wijesekera, Arjun Baokar, Lynn Tsai, Joel Reardon 等S&P 2017 · 被引用 156 次
- When Personalization Harms Performance: Reconsidering the Use of Group Attributes in PredictionVinith Menon Suriyakumar, Marzyeh Ghassemi, Berk UstunICML 2023 · 被引用 10 次
