Auditing Black-Box Prediction Models for Data Minimization Compliance
Bashir Rastegarpanah, Krishna P. Gummadi, Mark Crovella
摘要
In this paper, we focus on auditing black-box prediction models for compliance with the GDPR's data minimization principle. This principle restricts prediction models to use the minimal information that is necessary for performing the task at hand. Given the challenge of the black-box setting, our key idea is to check if each of the prediction model's input features is individually necessary by assigning it some constant value (i.e., applying a simple imputation) across all prediction instances, and measuring the extent to which the model outcomes would change. We introduce a metric for data minimization that is based on model instability under simple imputations. We extend the applicability of this metric from a finite sample model to a distributional setting by introducing a probabilistic data minimization guarantee, which we derive using a Bayesian approach. Furthermore, we address the auditing problem under a constraint on the number of queries to the prediction system. We formulate the problem of allocating a budget of system queries to feasible simple imputations (for investigating model instability) as a multi-armed bandit framework with probabilistic success metrics. We define two bandit problems for providing a probabilistic data minimization guarantee at a given confidence level: a decision problem given a data minimization level, and a measurement problem given a fixed query budget. We design efficient algorithms for these auditing problems using novel exploration strategies that expand classical bandit strategies. Our experiments with real-world prediction systems show that our auditing algorithms significantly outperform simpler benchmarks in both measurement and decision problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Active fairness auditingTom Yan, Chicheng ZhangICML 2022 · 被引用 34 次
- Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based ChatbotsJijie Zhou, Eryue Xu, Yaoyao Wu, Tianshi LiCHI 2025 · 被引用 15 次
- From Principle to Practice: Vertical Data Minimization for Machine LearningRobin Staab, Nikola Jovanovic, Mislav Balunovic, Martin T. VechevS&P 2024 · 被引用 10 次
- Data Minimization at Inference TimeCuong Tran, Ferdinando FiorettoNeurIPS 2023 · 被引用 8 次
它引用的顶会 Paper1
相关 Paper
- TeDA: A Testing Framework for Data Usage Auditing in Deep Learning Model DevelopmentXiangshan Gao, Jialuo Chen, Jingyi Wang, Jie Shi 等ISSTA 2024
- Online Fairness Auditing through Iterative RefinementPranav Maneriker, Codi Burley, Srinivasan ParthasarathyKDD 2023 · 被引用 6 次
- How Android Apps Break the Data Minimization Principle: An Empirical StudyShaokun Zhang, Hanwen Lei, Yuanpeng Wang, Ding Li 等ASE 2023 · 被引用 3 次
- "I'm not convinced that they don't collect more than is necessary": User-Controlled Data Minimization Design in Search EnginesTanusree Sharma, Lin Kyi, Yang Wang, Asia J. BiegaUSENIX Security 2024 · 被引用 5 次
- PAC-Private AlgorithmsMayuri Sridhar, Hanshen Xiao, Srinivas DevadasS&P 2025
