Size-adaptive Hypothesis Testing for Fairness
Antonio Ferrara, Francesco Cozzi, Alan Perotti, André Panisson, Francesco Bonchi
Abstract
Determining whether an algorithmic decision-making system discriminates against a specific demographic typically involves comparing a single point estimate of a fairness metric against a predefined threshold. This practice is statistically brittle: it ignores sampling error and treats small demographic subgroups the same as large ones. The problem intensifies in intersectional analyses, where multiple sensitive attributes are considered jointly, giving rise to a larger number of smaller groups. As these groups become more granular, the data representing them becomes too sparse for reliable estimation, and fairness metrics yield excessively wide confidence intervals, precluding meaningful conclusions about potential unfair treatments. In this paper, we introduce a unified, size-adaptive, hypothesis-testing framework that turns fairness assessment into an evidence-based statistical decision. Our contribution is twofold. (i) For sufficiently large subgroups, we prove a Central-Limit result for the statistical parity difference, leading to analytic confidence intervals and a Wald test whose type-I (false positive) error is guaranteed at level . (ii) For the long tail of small intersectional groups, we derive a fully Bayesian Dirichlet-multinomial estimator; Monte-Carlo credible intervals are calibrated for any sample size and naturally converge to Wald intervals as more data becomes available. We validate our approach empirically on benchmark datasets, demonstrating how our tests provide interpretable, statistically rigorous decisions under varying degrees of data availability and intersectionality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dae6051-68c3-4e21-9df3-8c134fcdf13dBuilds on3
- Maxmin-Fair Ranking: Individual Fairness under Group-Fairness ConstraintsDavid García-Soriano, Francesco BonchiKDD 2021 · 30 citations
- Bounding and Approximating Intersectional Fairness through Marginal FairnessMathieu Molina, Patrick LoiseauNeurIPS 2022 · 16 citations
- Beyond Shortest Paths: Node Fairness in Route RecommendationAntonio Ferrara, David García-Soriano, Francesco BonchiVLDB 2025 · 1 citation
Related papers
- Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian InferenceDisi Ji, Padhraic Smyth, Mark SteyversNeurIPS 2020 · 57 citations
- Evaluating model performance under worst-case subpopulationsMike Li, Hongseok Namkoong, Shangzhou XiaNeurIPS 2021 · 19 citations
- Monitoring Algorithmic FairnessThomas A. Henzinger, Mahyar Karimi, Konstantin Kueffner, Kaushik MallikCAV 2023 · 13 citations
- Bayes-Optimal Fair Classification with Multiple Sensitive FeaturesYi Yang, Yinghui Huang, Xiangyu ChangAAAI 2026 · 2 citations
- A Sequentially Fair Mechanism for Multiple Sensitive AttributesFrançois Hu, Philipp Ratz, Arthur CharpentierAAAI 2024 · 11 citations
