Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference
Disi Ji, Padhraic Smyth, Mark Steyvers
Abstract
We investigate the problem of reliably assessing group fairness when labeled examples are few but unlabeled examples are plentiful. We propose a general Bayesian framework that can augment labeled data with unlabeled data to produce more accurate and lower-variance estimates compared to methods based on labeled data alone. Our approach estimates calibrated scores for unlabeled examples in each group using a hierarchical latent variable model conditioned on labeled examples. This in turn allows for inference of posterior distributions with associated notions of uncertainty for a variety of group fairness metrics. We demonstrate that our approach leads to significant and consistent reductions in estimation error across multiple well-known fairness datasets, sensitive attributes, and predictive models. The results show the benefits of using both unlabeled data and Bayesian inference in terms of assessing whether a prediction model is fair or not.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47f8a0a8-451b-4d8f-8dde-c15f99185451Cited by top-tier papers8
- Scalable and Stable Surrogates for Flexible Classifiers with Fairness ConstraintsHarry Bendekgey, Erik B. SudderthNeurIPS 2021 · 23 citations
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 20 citations
- Bounding and Approximating Intersectional Fairness through Marginal FairnessMathieu Molina, Patrick LoiseauNeurIPS 2022 · 16 citations
- Access Denied: Meaningful Data Access for Quantitative Algorithm AuditsJuliette Zaccour, Reuben Binns, Luc RocherCHI 2025 · 9 citations
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag et al.NeurIPS 2025 · 9 citations
Related papers
- A Fair Bayesian Inference through Matched Gibbs PosteriorJihu Lee, Kunwoong Kim, Sehyun Park, Insung Kong et al.ICLR 2026
- Multiaccuracy and Multicalibration via Proxy GroupsBeepul Bharti, Mary Versa Clemens-Sewall, Paul H. Yi, Jeremias SulamICML 2025
- Size-adaptive Hypothesis Testing for FairnessAntonio Ferrara, Francesco Cozzi, Alan Perotti, André Panisson et al.NeurIPS 2025 · 2 citations
- Who's the (Multi-)Fairest of Them All: Rethinking Interpolation-Based Data Augmentation Through the Lens of MulticalibrationKarina Halevy, Karly Hou, Charumathi BadrinathAAAI 2025 · 2 citations
- Post-hoc bias scoring is optimal for fair classificationWenlong Chen, Yegor Klochkov, Yang LiuICLR 2024 · 12 citations
