Weak Proxies are Sufficient and Preferable for Fairness with Missing Sensitive Attributes
Zhaowei Zhu, Yuanshun Yao, Jiankai Sun, Hang Li, Yang Liu
Abstract
Evaluating fairness can be challenging in practice because the sensitive attributes of data are often inaccessible due to privacy constraints. The go-to approach that the industry frequently adopts is using off-the-shelf proxy models to predict the missing sensitive attributes, e.g. Meta [Alao et al., 2021] and Twitter [Belli et al., 2022] . Despite its popularity, there are three important questions unanswered: (1) Is directly using proxies efficacious in measuring fairness? (2) If not, is it possible to accurately evaluate fairness using proxies only? (3) Given the ethical controversy over inferring user private information, is it possible to only use weak (i.e. inaccurate) proxies in order to protect privacy? Our theoretical analyses show that directly using proxy models can give a false sense of (un)fairness. Second, we develop an algorithm that is able to measure fairness (provably) accurately with only three properly identified proxies. Third, we show that our algorithm allows the use of only weak proxies (e.g. with only 68.85% accuracy on COMPAS), adding an extra layer of protection on user privacy. Experiments validate our theoretical analyses and show our algorithm can effectively measure and mitigate bias. Our results imply a set of practical guidelines for practitioners on how to use proxies properly. Code is available at github.com/UCSC-REAL/fair-eval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9071f136-1fd6-4539-a644-eceb714d222bCited by top-tier papers10
- Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language ModelsZhaowei Zhu, Jialu Wang, Hao Cheng, Yang LiuICLR 2024 · 30 citations
- Fairness without Harm: An Influence-Guided Active Sampling ApproachJinlong Pang, Jialu Wang, Zhaowei Zhu, Yuanshun Yao et al.NeurIPS 2024 · 13 citations
- Post-hoc bias scoring is optimal for fair classificationWenlong Chen, Yegor Klochkov, Yang LiuICLR 2024 · 12 citations
- On the Maximal Local Disparity of Fairness-Aware ClassifiersJinqiu Jin, Haoxuan Li, Fuli FengICML 2024 · 5 citations
- Distributionally Generative Augmentation for Fair Facial Attribute ClassificationFengda Zhang, Qianpei He, Kun Kuang, Jiashuo Liu et al.CVPR 2024 · 4 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 280 citations
- Deep Learning with Label Differential PrivacyBadih Ghazi, Noah Golowich, Ravi Kumar, Pasin Manurangsi et al.NeurIPS 2021 · 193 citations
- Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for SupportMichael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan et al.CSCW 2022 · 149 citations
Related papers
- Multiaccuracy and Multicalibration via Proxy GroupsBeepul Bharti, Mary Versa Clemens-Sewall, Paul H. Yi, Jeremias SulamICML 2025
- Assessing Fairness in the Presence of Missing DataYiliang Zhang, Qi LongNeurIPS 2021 · 51 citations
- Use Privacy in Data-Driven Systems: Theory and Experiments with Machine Learnt ProgramsAnupam Datta, Matthew Fredrikson, Gihyuk Ko, Piotr Mardziel et al.CCS 2017 · 63 citations
- FairLISA: Fair User Modeling with Limited Sensitive Attributes InformationZheng Zhang, Qi Liu, Hao Jiang, Fei Wang et al.NeurIPS 2023 · 42 citations
- Fair Data Pre-Processing with Imperfect Attribute SpaceYing Zheng, Yangfan Jiang, Kian-Lee TanSIGMOD 2026
