How Accurate Does It Feel? - Human Perception of Different Types of Classification Mistakes
Andrea Papenmeier, Dagmar Kern, Daniel Hienert, Yvonne Kammerer, Christin Seifert
摘要
Supervised machine learning utilizes large datasets, often with ground truth labels annotated by humans. While some data points are easy to classify, others are hard to classify, which reduces the inter-annotator agreement. This causes noise for the classifier and might affect the user's perception of the classifier's performance. In our research, we investigated whether the classification difficulty of a data point influences how strongly a prediction mistake reduces the "perceived accuracy". In an experimental online study, 225 participants interacted with three fictive classifiers with equal accuracy (73%). The classifiers made prediction mistakes on three different types of data points (easy, difficult, impossible). After the interaction, participants judged the classifier's accuracy. We found that not all prediction mistakes reduced the perceived accuracy equally. Furthermore, the perceived accuracy differed significantly from the calculated accuracy. To conclude, accuracy and related measures seem unsuitable to represent how users perceive the performance of classifiers.
• Human-centered computing → Empirical studies in HCI; Human computer interaction (HCI).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Nurturing Capabilities: Unpacking the Gap in Human-Centered Evaluations of AI-Based SystemsAman Khullar, Nikhil Nalin, Abhishek Prasad, Ann John Mampilli 等CHI 2025 · 被引用 14 次
- Everybody's Got ML, Tell Me What Else You Have: Practitioners' Perception of ML-Based Security Tools and ExplanationsJaron Mink, Hadjer Benkraouda, Limin Yang, Arridhana Ciptadi 等S&P 2023
它引用的顶会 Paper14
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic RetinopathyEmma Beede, Elizabeth Elliott Baylor, Fred Hersch, Anna Iurchenko 等CHI 2020 · 被引用 589 次
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 被引用 428 次
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
相关 Paper
- Human and AI Perceptual Differences in Image Classification ErrorsMinghao Liu, Jiaheng Wei, Yang Liu, James DavisAAAI 2025 · 被引用 11 次
- Truth or Dare: Understanding and Predicting How Users Lie and Provide Untruthful Data OnlineKopo M. Ramokapane, Gaurav Misra, Jose M. Such, Sören PreibuschCHI 2021 · 被引用 16 次
- When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning ModelsAmy Rechkemmer, Ming YinCHI 2022 · 被引用 94 次
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag 等NeurIPS 2025 · 被引用 9 次
- It's Trying Too Hard To Look Real: Deepfake Moderation Mistakes and Identity-Based BiasJaron Mink, Miranda Wei, Collins W. Munyendo, Kurt Hugenberg 等CHI 2024 · 被引用 10 次
