When AI Difficulty Is Easy: The Explanatory Power of Predicting IRT Difficulty
Fernando Martínez-Plumed, David Castellano Falcón, Carlos Monserrat Aranda, José Hernández-Orallo
Abstract
One of challenges of artificial intelligence as a whole is robustness. Many issues such as adversarial examples, out of distribution performance, Clever Hans phenomena, and the wider areas of AI evaluation and explainable AI, have to do with the following question: Did the system fail because it is a hard instance or because something else? In this paper we address this question with a generic method for estimating IRT-based instance difficulty for a wide range of AI domains covering several areas, from supervised feature-based classification to automated reasoning. We show how to estimate difficulty systematically using off-the-shelf machine learning regression models. We illustrate the usefulness of this estimation for a range of applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 656ae730-5e9e-4d89-950a-401584c7131aCited by top-tier papers1
Ask how each one uses itRelated papers
- Transferability and Hardness of Supervised Classification TasksAnh Tuan Tran, Cuong V. Nguyen, Tal HassnerICCV 2019 · 201 citations
- The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden RepresentationsYubo Zhu, Dongrui Liu, Zecheng Lin, Wei Tong et al.EMNLP 2025 · 1 citation
- Reliable and Efficient Amortized Model-based EvaluationSang T. Truong, Yuheng Tu, Percy Liang, Bo Li et al.ICML 2025
- ILDAE: Instance-Level Difficulty Analysis of Evaluation DataNeeraj Varshney, Swaroop Mishra, Chitta BaralACL 2022 · 21 citations
- Humanly Certifying Superhuman ClassifiersQiongkai Xu, Christian Walder, Chenchen XuICLR 2023
