A Multimodal Automated Interpretability Agent
Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, Antonio Torralba
摘要
This paper describes MAIA, a Multimodal Automated Interpretability Agent. MAIA is a system that uses neural models to automate neural model understanding tasks like feature interpretation and failure mode discovery. It equips a pre-trained vision-language model with a set of tools that support iterative experimentation on subcomponents of other models to explain their behavior. These include tools commonly used by human interpretability researchers: for synthesizing and editing inputs, computing maximally activating exemplars from real-world datasets, and summarizing and describing experimental results. Interpretability experiments proposed by MAIA compose these tools to describe and explain system behavior. We evaluate applications of MAIA to computer vision models. We first characterize MAIA's ability to describe (neuron-level) features in learned representations of images. Across several trained models and a novel dataset of synthetic vision neurons with paired ground-truth descriptions, MAIA produces descriptions comparable to those generated by expert human experimenters. We then show that MAIA can aid in two additional interpretability tasks: reducing sensitivity to spurious features, and automatically identifying inputs likely to be mis-classified. †
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Narrow Finetuning Leaves Clearly Readable Traces in Activation DifferencesJulian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt 等ICLR 2026 · 被引用 29 次
- Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski GeometryThomas Fel, Binxu Wang, Michael A. Lepori, Matthew Kowal 等ICLR 2026 · 被引用 28 次
- Learning Concept Bottleneck Models from Mechanistic ExplanationsAntonio De Santis, Schrasing Tong, Marco Brambilla, Lalana KagalICLR 2026 · 被引用 6 次
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured LanguageAngie W. Boggust, Donghao Ren, Yannick Assogba, Dominik Moritz 等ICLR 2026 · 被引用 3 次
- Language Models Can Explain Visual Features via SteeringJavier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres, Daniel Hinjos 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
相关 Paper
- HiBug: On Human-Interpretable Model DebugMuxi Chen, Yu Li, Qiang XuNeurIPS 2023 · 被引用 22 次
- RetouchAgent: Towards Interactive and Explainable Image Retouching with MLLM AgentsShuo Zhang, Xinyu YangAAAI 2026
- Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency AnalysisEldon Schoop, Xin Zhou, Gang Li, Zhourong Chen 等CHI 2022 · 被引用 29 次
- Automated Detection of Visual Attribute Reliance with a Self-Reflective AgentChristy Li, Josep López Camuñas, Jake Thomas Touchet, Jacob Andreas 等NeurIPS 2025 · 被引用 2 次
- Red Teaming Deep Neural Networks with Feature Synthesis ToolsStephen Casper, Tong Bu, Yuxiao Li, Jiawei Li 等NeurIPS 2023 · 被引用 23 次
