Fairness Evaluation with Item Response Theory
Ziqi Xu, Sevvandi Kandanaarachchi, Cheng Soon Ong, Eirini Ntoutsi
Abstract
Item Response Theory (IRT) has been widely used in educational psychometrics to assess student ability, as well as the difficulty and discrimination of test questions. In this context, discrimination specifically refers to how effectively a question distinguishes between students of different ability levels, and it does not carry any connotation related to fairness. In recent years, IRT has been successfully used to evaluate the predictive performance of Machine Learning (ML) models, but this paper marks its first application in fairness evaluation. In this paper, we propose a novel Fair-IRT framework to evaluate a set of predictive models on a set of individuals, while simultaneously eliciting specific parameters, namely, the ability to make fair predictions (a feature of predictive models), as well as the discrimination and difficulty of individuals that affect the prediction results. Furthermore, we conduct a series of experiments to comprehensively understand the implications of these parameters for fairness evaluation. Detailed explanations for item characteristic curves (ICCs) are provided for particular individuals. We propose the flatness of ICCs to disentangle the unfairness between individuals and predictive models. The experiments demonstrate the effectiveness of this framework as a fairness evaluation tool. Two real-world case studies illustrate its potential application in evaluating fairness in both classification and regression tasks. Our paper aligns well with the Responsible Web track by proposing a Fair-IRT framework to evaluate fairness in ML models, which directly contributes to the development of a more inclusive, equitable, and trustworthy AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1951ed4a-68d9-4bfb-9e68-745e6e3a52c4Cited by top-tier papers2
- FairGE: Fairness-Aware Graph Encoding in Incomplete Social NetworksRenqiang Luo, Huafei Huang, Tao Tang, Jing Ren et al.WWW 2026 · 1 citation
- Noise-Aware Graph-Based Cognitive Diagnostic Framework Through Low-Rank AlignmentGuixian Zhang, Yanmei Zhang, Guan Yuan, Shang Liu et al.AAAI 2026
Builds on2
Related papers
- Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling EstimationSang Truong, Yuheng Tu, Rylan Schaeffer, Sanmi KoyejoICML 2026
- Approximation-guided Fairness Testing through Discriminatory Space AnalysisZhenjiang Zhao, Takahisa Toda, Takashi KitamuraASE 2024
- Achieving Equalized Odds by Resampling Sensitive AttributesYaniv Romano, Stephen Bates, Emmanuel J. CandèsNeurIPS 2020 · 65 citations
- Evaluating Fairness Using Permutation TestsCyrus DiCiccio, Sriram Vasudevan, Kinjal Basu, Krishnaram Kenthapadi et al.KDD 2020 · 2 citations
- Efficient white-box fairness testing through gradient searchLingfeng Zhang, Yueling Zhang, Min ZhangISSTA 2021 · 51 citations
