Designing Alternative Representations of Confusion Matrices to Support Non-Expert Public Understanding of Algorithm Performance
Hong Shen, Haojian Jin, Ángel Alexander Cabrera, Adam Perer, Haiyi Zhu, Jason I. Hong
Abstract
Ensuring effective public understanding of algorithmic decisions that are powered by machine learning techniques has become an urgent task with the increasing deployment of AI systems into our society. In this work, we present a concrete step toward this goal by redesigning confusion matrices for binary classification to support non-experts in understanding the performance of machine learning models. Through interviews (n=7) and a survey (n=102), we mapped out two major sets of challenges lay people have in understanding standard confusion matrices: the general terminologies and the matrix design. We further identified three sub-challenges regarding the matrix design, namely, confusion about the direction of reading the data, layered relations and quantities involved. We then conducted an online experiment with 483 participants to evaluate how effective a series of alternative representations target each of those challenges in the context of an algorithm for making recidivism predictions. We developed three levels of questions to evaluate users' objective understanding. We assessed the effectiveness of our alternatives for accuracy in answering those questions, completion time, and subjective understanding. Our results suggest that (1) only by contextualizing terminologies can we significantly improve users' understanding and (2) flow charts, which help point out the direction of reading the data, were most useful in improving objective understanding. Our findings set the stage for developing more intuitive and generally understandable representations of the performance of machine learning models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 933f075a-c00c-4e46-b4c7-3e923abef8efCited by top-tier papers14
- Seeing Like a Toolkit: How Toolkits Envision the Work of AI EthicsRichmond Y. Wong, Michael A. Madaio, Nick MerrillCSCW 2023 · 108 citations
- Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis DiagnosisShao Zhang, Jianing Yu, Xuhai Xu, Changchang Yin et al.CHI 2024 · 95 citations
- Effect of Information Presentation on Fairness Perceptions of Machine Learning PredictorsNiels van Berkel, Jorge Gonçalves, Daniel Russo, Simo Hosio et al.CHI 2021 · 84 citations
- Deliberating with AI: Improving Decision-Making for the Future through Participatory AI Design and Stakeholder DeliberationAngie Zhang, Olympia Walker, Kaci Nguyen, Jiajun Dai et al.CSCW 2023 · 71 citations
- Neo: Generalizing Confusion Matrix Visualization to Hierarchical and Multi-Output LabelsJochen Görtler, Fred Hohman, Dominik Moritz, Kanit Wongsuphasawat et al.CHI 2022 · 70 citations
Related papers
- Never mind the metrics - what about the uncertainty? Visualising binary confusion matrix metric distributions to put performance in perspectiveDavid R. Lovell, Dimity Miller, Jaiden Capra, Andrew P. BradleyICML 2023 · 3 citations
- My Model is Unfair, Do People Even Care? Visual Design Affects Trust and Perceived Bias in Machine LearningAimen Gaba, Zhanna Kaufman, Jason Cheung, Marie Shvakel et al.IEEE VIS 2023 · 20 citations
- The Impact of Algorithmic Risk Assessments on Human Predictions and its Analysis via Crowdsourcing StudiesRiccardo Fogliato, Alexandra Chouldechova, Zachary C. LiptonCSCW 2021 · 25 citations
- Evaluating the Interpretability of Generative Models by Interactive ReconstructionAndrew Slavin Ross, Nina Chen, Elisa Zhao Hang, Elena L. Glassman et al.CHI 2021 · 40 citations
- Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language ModelsMarvin Pafla, Kate Larson, Mark HancockCHI 2024 · 16 citations
