Fairness Evaluation with Item Response Theory
Ziqi Xu, Sevvandi Kandanaarachchi, Cheng Soon Ong, Eirini Ntoutsi
摘要
Item Response Theory (IRT) has been widely used in educational psychometrics to assess student ability, as well as the difficulty and discrimination of test questions. In this context, discrimination specifically refers to how effectively a question distinguishes between students of different ability levels, and it does not carry any connotation related to fairness. In recent years, IRT has been successfully used to evaluate the predictive performance of Machine Learning (ML) models, but this paper marks its first application in fairness evaluation. In this paper, we propose a novel Fair-IRT framework to evaluate a set of predictive models on a set of individuals, while simultaneously eliciting specific parameters, namely, the ability to make fair predictions (a feature of predictive models), as well as the discrimination and difficulty of individuals that affect the prediction results. Furthermore, we conduct a series of experiments to comprehensively understand the implications of these parameters for fairness evaluation. Detailed explanations for item characteristic curves (ICCs) are provided for particular individuals. We propose the flatness of ICCs to disentangle the unfairness between individuals and predictive models. The experiments demonstrate the effectiveness of this framework as a fairness evaluation tool. Two real-world case studies illustrate its potential application in evaluating fairness in both classification and regression tasks. Our paper aligns well with the Responsible Web track by proposing a Fair-IRT framework to evaluate fairness in ML models, which directly contributes to the development of a more inclusive, equitable, and trustworthy AI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FairGE: Fairness-Aware Graph Encoding in Incomplete Social NetworksRenqiang Luo, Huafei Huang, Tao Tang, Jing Ren 等WWW 2026 · 被引用 1 次
- Noise-Aware Graph-Based Cognitive Diagnostic Framework Through Low-Rank AlignmentGuixian Zhang, Yanmei Zhang, Guan Yuan, Shang Liu 等AAAI 2026
它引用的顶会 Paper2
相关 Paper
- Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling EstimationSang Truong, Yuheng Tu, Rylan Schaeffer, Sanmi KoyejoICML 2026
- Approximation-guided Fairness Testing through Discriminatory Space AnalysisZhenjiang Zhao, Takahisa Toda, Takashi KitamuraASE 2024
- Achieving Equalized Odds by Resampling Sensitive AttributesYaniv Romano, Stephen Bates, Emmanuel J. CandèsNeurIPS 2020 · 被引用 65 次
- Evaluating Fairness Using Permutation TestsCyrus DiCiccio, Sriram Vasudevan, Kinjal Basu, Krishnaram Kenthapadi 等KDD 2020 · 被引用 2 次
- Efficient white-box fairness testing through gradient searchLingfeng Zhang, Yueling Zhang, Min ZhangISSTA 2021 · 被引用 51 次
