How Accurate Does It Feel? - Human Perception of Different Types of Classification Mistakes
Andrea Papenmeier, Dagmar Kern, Daniel Hienert, Yvonne Kammerer, Christin Seifert
Abstract
Supervised machine learning utilizes large datasets, often with ground truth labels annotated by humans. While some data points are easy to classify, others are hard to classify, which reduces the inter-annotator agreement. This causes noise for the classifier and might affect the user's perception of the classifier's performance. In our research, we investigated whether the classification difficulty of a data point influences how strongly a prediction mistake reduces the "perceived accuracy". In an experimental online study, 225 participants interacted with three fictive classifiers with equal accuracy (73%). The classifiers made prediction mistakes on three different types of data points (easy, difficult, impossible). After the interaction, participants judged the classifier's accuracy. We found that not all prediction mistakes reduced the perceived accuracy equally. Furthermore, the perceived accuracy differed significantly from the calculated accuracy. To conclude, accuracy and related measures seem unsuitable to represent how users perceive the performance of classifiers.
• Human-centered computing → Empirical studies in HCI; Human computer interaction (HCI).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Nurturing Capabilities: Unpacking the Gap in Human-Centered Evaluations of AI-Based SystemsAman Khullar, Nikhil Nalin, Abhishek Prasad, Ann John Mampilli et al.CHI 2025 · 14 citations
- Everybody's Got ML, Tell Me What Else You Have: Practitioners' Perception of ML-Based Security Tools and ExplanationsJaron Mink, Hadjer Benkraouda, Limin Yang, Arridhana Ciptadi et al.S&P 2023
Builds on14
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic RetinopathyEmma Beede, Elizabeth Elliott Baylor, Fred Hersch, Anna Iurchenko et al.CHI 2020 · 589 citations
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 428 citations
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
Related papers
- Human and AI Perceptual Differences in Image Classification ErrorsMinghao Liu, Jiaheng Wei, Yang Liu, James DavisAAAI 2025 · 11 citations
- Truth or Dare: Understanding and Predicting How Users Lie and Provide Untruthful Data OnlineKopo M. Ramokapane, Gaurav Misra, Jose M. Such, Sören PreibuschCHI 2021 · 16 citations
- When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning ModelsAmy Rechkemmer, Ming YinCHI 2022 · 94 citations
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag et al.NeurIPS 2025 · 9 citations
- It's Trying Too Hard To Look Real: Deepfake Moderation Mistakes and Identity-Based BiasJaron Mink, Miranda Wei, Collins W. Munyendo, Kurt Hugenberg et al.CHI 2024 · 10 citations
