On the Epistemic Limits of Personalized Prediction
Lucas Monteiro Paes, Carol Xuan Long, Berk Ustun, Flávio P. Calmon
摘要
Machine learning models are often personalized by using group attributes that encode personal characteristics (e.g., sex, age group, HIV status). In such settings, individuals expect to receive more accurate predictions in return for disclosing group attributes to the personalized model. We study when we can tell that a personalized model upholds this principle for every group who provides personal data. We introduce a metric called the benefit of personalization (BoP) to measure the smallest gain in accuracy that any group expects to receive from a personalized model. We describe how the BoP can be used to carry out basic routines to audit a personalized model, including: (i) hypothesis tests to check that a personalized model improves performance for every group; (ii) estimation procedures to bound the minimum gain in personalization. We characterize the reliability of these routines in a finite-sample regime and present minimax bounds on both the probability of error for BoP hypothesis tests and the mean-squared error of BoP estimates. Our results show that we can only claim that personalization improves performance for each group who provides data when we explicitly limit the number of group attributes used by a personalized model. In particular, we show that it is impossible to reliably verify that a personalized classifier with k ≥ 19 binary group attributes will benefit every group who provides personal data using a dataset of n = 8 × 10 9 samples -one for each person in the world. * Equal contribution. 36th Conference on Neural Information Processing Systems (NeurIPS 2022). TRAINING DATA AUDITING DATA Group n R(h p ) R(h 0 ) R(h 0 ) -R(h p ) n R(h p ) R(h 0 ) R(h 0 ) -R(h p ) Female, W, NR
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Direct Alignment with Heterogeneous PreferencesAli Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri 等NeurIPS 2025 · 被引用 26 次
- Allocation Requires Prediction Only if Inequality Is LowAli Shirali, Rediet Abebe, Moritz HardtICML 2024 · 被引用 12 次
- When Personalization Harms Performance: Reconsidering the Use of Group Attributes in PredictionVinith Menon Suriyakumar, Marzyeh Ghassemi, Berk UstunICML 2023 · 被引用 10 次
- Participatory Personalization in ClassificationHailey Joren, Chirag Nagpal, Katherine A. Heller, Berk UstunNeurIPS 2023 · 被引用 7 次
- No Free Delivery Service: Epistemic limits of passive data collection in complex social systemsMaximilian NickelNeurIPS 2024 · 被引用 1 次
它引用的顶会 Paper3
- Minimax Pareto Fairness: A Multi Objective PerspectiveNatalia Martínez, Martín Bertrán, Guillermo SapiroICML 2020 · 被引用 232 次
- Testing Group Fairness via Optimal Transport ProjectionsNian Si, Karthyek Murthy, Jose H. Blanchet, Viet Anh NguyenICML 2021 · 被引用 37 次
- Regulating algorithmic filtering on social mediaSarah Huiyi Cen, Devavrat ShahNeurIPS 2021 · 被引用 12 次
相关 Paper
- When Machine Learning Gets Personal: Evaluating Prediction and ExplanationLouisa Cornelis, Guillermo Bernardez, Haewon Jeong, Nina MiolaneICLR 2026
- Robust Optimization for Fairness with Noisy Protected GroupsSerena Lutong Wang, Wenshuo Guo, Harikrishna Narasimhan, Andrew Cotter 等NeurIPS 2020 · 被引用 134 次
- Auditing Black-Box Prediction Models for Data Minimization ComplianceBashir Rastegarpanah, Krishna P. Gummadi, Mark CrovellaNeurIPS 2021 · 被引用 24 次
- Access Denied: Meaningful Data Access for Quantitative Algorithm AuditsJuliette Zaccour, Reuben Binns, Luc RocherCHI 2025 · 被引用 9 次
- Beyond Membership: Limitations of Add/Remove Adjacency in Differential PrivacyGauri Pradhan, Joonas Jälkö, Santiago Zanella-Béguelin, Antti HonkelaICLR 2026
