Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior
Angie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik Strobelt
Abstract
Saliency methods — techniques to identify the importance of input features on a model’s output — are a common step in understanding neural network behavior. However, interpreting saliency requires tedious manual inspection to identify and aggregate patterns in model behavior, resulting in ad hoc or cherry-picked analysis. To address these concerns, we present Shared Interest: metrics for comparing model reasoning (via saliency) to human reasoning (via ground truth annotations). By providing quantitative descriptors, Shared Interest enables ranking, sorting, and aggregating inputs, thereby facilitating large-scale systematic analysis of model behavior. We use Shared Interest to identify eight recurring patterns in model behavior, such as cases where contextual features or a subset of ground truth features are most important to the model. Working with representative real-world users, we show how Shared Interest can be used to decide if a model is trustworthy, uncover issues missed in manual analyses, and enable interactive probing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-MakingShuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng et al.CHI 2025 · 113 citations
- Selective Explanations: Leveraging Human Input to Align Explainable AIVivian Lai, Yiming Zhang, Chacha Chen, Q. Vera Liao et al.CSCW 2023 · 49 citations
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt et al.ICCV 2025 · 47 citations
- Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision MakingSara Salimzadeh, Gaole He, Ujwal GadirajuCHI 2024 · 40 citations
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu et al.IEEE VIS 2024 · 23 citations
Builds on5
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 451 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram et al.AAAI 2020 · 204 citations
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas et al.ICML 2020 · 146 citations
- Overinterpretation reveals image classification model pathologiesBrandon Carter, Siddhartha Jain, Jonas Mueller, David GiffordNeurIPS 2021 · 59 citations
Related papers
- Sanity Simulations for Saliency MethodsJoon Sik Kim, Gregory Plumb, Ameet TalwalkarICML 2022 · 24 citations
- New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and SoundArushi Gupta, Nikunj Saunshi, Dingli Yu, Kaifeng Lyu et al.NeurIPS 2022 · 12 citations
- Passive attention in artificial neural networks predicts human visual selectivityThomas A. Langlois, H. Charles Zhao, Erin Grant, Ishita Dasgupta et al.NeurIPS 2021 · 19 citations
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 167 citations
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 158 citations
