Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior
Angie W. Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik Strobelt
摘要
Saliency methods — techniques to identify the importance of input features on a model’s output — are a common step in understanding neural network behavior. However, interpreting saliency requires tedious manual inspection to identify and aggregate patterns in model behavior, resulting in ad hoc or cherry-picked analysis. To address these concerns, we present Shared Interest: metrics for comparing model reasoning (via saliency) to human reasoning (via ground truth annotations). By providing quantitative descriptors, Shared Interest enables ranking, sorting, and aggregating inputs, thereby facilitating large-scale systematic analysis of model behavior. We use Shared Interest to identify eight recurring patterns in model behavior, such as cases where contextual features or a subset of ground truth features are most important to the model. Working with representative real-world users, we show how Shared Interest can be used to decide if a model is trustworthy, uncover issues missed in manual analyses, and enable interactive probing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-MakingShuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng 等CHI 2025 · 被引用 113 次
- Selective Explanations: Leveraging Human Input to Align Explainable AIVivian Lai, Yiming Zhang, Chacha Chen, Q. Vera Liao 等CSCW 2023 · 被引用 49 次
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt 等ICCV 2025 · 被引用 47 次
- Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision MakingSara Salimzadeh, Gaole He, Ujwal GadirajuCHI 2024 · 被引用 40 次
- LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language ModelsMinsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu 等IEEE VIS 2024 · 被引用 23 次
它引用的顶会 Paper5
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 被引用 451 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- Sanity Checks for Saliency MetricsRichard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram 等AAAI 2020 · 被引用 204 次
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas 等ICML 2020 · 被引用 146 次
- Overinterpretation reveals image classification model pathologiesBrandon Carter, Siddhartha Jain, Jonas Mueller, David GiffordNeurIPS 2021 · 被引用 59 次
相关 Paper
- Sanity Simulations for Saliency MethodsJoon Sik Kim, Gregory Plumb, Ameet TalwalkarICML 2022 · 被引用 24 次
- New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and SoundArushi Gupta, Nikunj Saunshi, Dingli Yu, Kaifeng Lyu 等NeurIPS 2022 · 被引用 12 次
- Passive attention in artificial neural networks predicts human visual selectivityThomas A. Langlois, H. Charles Zhao, Erin Grant, Ishita Dasgupta 等NeurIPS 2021 · 被引用 19 次
- Do Feature Attribution Methods Correctly Attribute Features?Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie ShahAAAI 2022 · 被引用 167 次
- A Diagnostic Study of Explainability Techniques for Text ClassificationPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinEMNLP 2020 · 被引用 158 次
