Weak Proxies are Sufficient and Preferable for Fairness with Missing Sensitive Attributes
Zhaowei Zhu, Yuanshun Yao, Jiankai Sun, Hang Li, Yang Liu
摘要
Evaluating fairness can be challenging in practice because the sensitive attributes of data are often inaccessible due to privacy constraints. The go-to approach that the industry frequently adopts is using off-the-shelf proxy models to predict the missing sensitive attributes, e.g. Meta [Alao et al., 2021] and Twitter [Belli et al., 2022] . Despite its popularity, there are three important questions unanswered: (1) Is directly using proxies efficacious in measuring fairness? (2) If not, is it possible to accurately evaluate fairness using proxies only? (3) Given the ethical controversy over inferring user private information, is it possible to only use weak (i.e. inaccurate) proxies in order to protect privacy? Our theoretical analyses show that directly using proxy models can give a false sense of (un)fairness. Second, we develop an algorithm that is able to measure fairness (provably) accurately with only three properly identified proxies. Third, we show that our algorithm allows the use of only weak proxies (e.g. with only 68.85% accuracy on COMPAS), adding an extra layer of protection on user privacy. Experiments validate our theoretical analyses and show our algorithm can effectively measure and mitigate bias. Our results imply a set of practical guidelines for practitioners on how to use proxies properly. Code is available at github.com/UCSC-REAL/fair-eval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language ModelsZhaowei Zhu, Jialu Wang, Hao Cheng, Yang LiuICLR 2024 · 被引用 30 次
- Fairness without Harm: An Influence-Guided Active Sampling ApproachJinlong Pang, Jialu Wang, Zhaowei Zhu, Yuanshun Yao 等NeurIPS 2024 · 被引用 13 次
- Post-hoc bias scoring is optimal for fair classificationWenlong Chen, Yegor Klochkov, Yang LiuICLR 2024 · 被引用 12 次
- On the Maximal Local Disparity of Fairness-Aware ClassifiersJinqiu Jin, Haoxuan Li, Fuli FengICML 2024 · 被引用 5 次
- Distributionally Generative Augmentation for Fair Facial Attribute ClassificationFengda Zhang, Qianpei He, Kun Kuang, Jiashuo Liu 等CVPR 2024 · 被引用 4 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 被引用 280 次
- Deep Learning with Label Differential PrivacyBadih Ghazi, Noah Golowich, Ravi Kumar, Pasin Manurangsi 等NeurIPS 2021 · 被引用 193 次
- Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for SupportMichael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan 等CSCW 2022 · 被引用 149 次
相关 Paper
- Multiaccuracy and Multicalibration via Proxy GroupsBeepul Bharti, Mary Versa Clemens-Sewall, Paul H. Yi, Jeremias SulamICML 2025
- Assessing Fairness in the Presence of Missing DataYiliang Zhang, Qi LongNeurIPS 2021 · 被引用 51 次
- Use Privacy in Data-Driven Systems: Theory and Experiments with Machine Learnt ProgramsAnupam Datta, Matthew Fredrikson, Gihyuk Ko, Piotr Mardziel 等CCS 2017 · 被引用 63 次
- FairLISA: Fair User Modeling with Limited Sensitive Attributes InformationZheng Zhang, Qi Liu, Hao Jiang, Fei Wang 等NeurIPS 2023 · 被引用 42 次
- Fair Data Pre-Processing with Imperfect Attribute SpaceYing Zheng, Yangfan Jiang, Kian-Lee TanSIGMOD 2026
