Manipulating and Measuring Model Interpretability
Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, Hanna M. Wallach
Abstract
With machine learning models being increasingly used to aid decision making even in high-stakes domains, there has been a growing interest in developing interpretable models. Although many supposedly interpretable models have been proposed, there have been relatively few experimental studies investigating whether these models achieve their intended effects, such as making people more closely follow a model’s predictions when it is beneficial for them to do so or enabling them to detect when a model has made a mistake. We present a sequence of pre-registered experiments (N = 3, 800) in which we showed participants functionally identical models that varied only in two factors commonly thought to make machine learning models more or less interpretable: the number of features and the transparency of the model (i.e., whether the model internals are clear or black box). Predictably, participants who saw a clear model with few features could better simulate the model’s predictions. However, we did not find that participants more closely followed its predictions. Furthermore, showing participants a clear model meant that they were less able to detect and correct for the model’s sizable mistakes, seemingly due to information overload. These counterintuitive findings emphasize the importance of testing over intuition when developing interpretable models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9958c0f7-cca8-44f2-bf60-64d329fa6348Cited by top-tier papers123
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana et al.CHI 2020 · 541 citations
- Expanding Explainability: Towards Social Transparency in AI systemsUpol Ehsan, Q. Vera Liao, Michael J. Muller, Mark O. Riedl et al.CHI 2021 · 505 citations
- Problems with Shapley-value-based explanations as feature importance measuresI. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, Sorelle A. FriedlerICML 2020 · 458 citations
Builds on3
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana et al.CHI 2020 · 541 citations
- COGAM: Measuring and Moderating Cognitive Load in Machine Learning Model ExplanationsAshraf M. Abdul, Christian von der Weth, Mohan S. Kankanhalli, Brian Y. LimCHI 2020 · 92 citations
Related papers
- Interpretability Gone Bad: The Role of Bounded Rationality in How Practitioners Understand Machine LearningHarmanpreet Kaur, Matthew R. Conrad, Davis Rule, Cliff Lampe et al.CSCW 2024 · 14 citations
- Impact of Model Interpretability and Outcome Feedback on Trust in AIDaehwan Ahn, Abdullah Almaatouq, Monisha Gulabani, Kartik HosanagarCHI 2024 · 33 citations
- Does Explainable Artificial Intelligence Improve Human Decision-Making?Yasmeen Alufaisan, Laura R. Marusich, Jonathan Z. Bakdash, Yan Zhou et al.AAAI 2021 · 135 citations
- When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning ModelsAmy Rechkemmer, Ming YinCHI 2022 · 94 citations
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 3 citations
