FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI Methods
Robin Hesse, Simone Schaub-Meyer, Stefan Roth
Abstract
The field of explainable artificial intelligence (XAI) aims to uncover the inner workings of complex deep neural models. While being crucial for safety-critical domains, XAI inherently lacks ground-truth explanations, making its automatic evaluation an unsolved problem. We address this challenge by proposing a novel synthetic vision dataset, named FunnyBirds, and accompanying automatic evaluation protocols. Our dataset allows performing semantically meaningful image interventions, e.g., removing individual object parts, which has three important implications. First, it enables analyzing explanations on a part level, which is closer to human comprehension than existing methods that evaluate on a pixel level. Second, by comparing the model output for inputs with removed parts, we can estimate ground-truth part importances that should be reflected in the explanations. Third, by mapping individual explanations into a common space of part importances, we can analyze a variety of different explanation types in a single common framework. Using our tools, we report results for 24 different combinations of neural models and XAI methods, demonstrating the strengths and weaknesses of the assessed methods in a fully automatic and systematic manner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityWalid Bousselham, Angie W. Boggust, Sofian Chaybouti, Hendrik Strobelt et al.ICCV 2025 · 47 citations
- Red Teaming Deep Neural Networks with Feature Synthesis ToolsStephen Casper, Tong Bu, Yuxiao Li, Jiawei Li et al.NeurIPS 2023 · 23 citations
- SUB: Benchmarking CBM Generalization via Synthetic Attribute SubstitutionsJessica Bader, Leander Girrbach, Stephan Alaniz, Zeynep AkataICCV 2025 · 8 citations
- EPIC: Explanation of Pretrained Image Classification Networks via PrototypesPiotr Borycki, Magdalena Tredowicz, Szymon Janusz, Jacek Tabor et al.AAAI 2026 · 4 citations
- Soft Local Completeness: Rethinking Completeness in XAIZiv Weiss Haddad, Oren Barkan, Yehonatan Elisha, Noam KoenigsteinICCV 2025 · 2 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability MethodsJulien Colin, Thomas Fel, Rémi Cadène, Thomas SerreNeurIPS 2022 · 147 citations
- The effectiveness of feature attribution methods and its correlation with automatic evaluation scoresGiang Nguyen, Daeyoung Kim, Anh NguyenNeurIPS 2021 · 128 citations
Related papers
- SIC: Similarity-Based Interpretable Image Classification with Neural NetworksTom Nuno Wolf, Emre Kavak, Fabian Bongratz, Christian WachingerICCV 2025 · 1 citation
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 62 citations
- Right for the Right Concept: Revising Neuro-Symbolic Concepts by Interacting With Their ExplanationsWolfgang Stammer, Patrick Schramowski, Kristian KerstingCVPR 2021
- Explain Any Concept: Segment Anything Meets Concept-Based ExplanationAo Sun, Pingchuan Ma, Yuanyuan Yuan, Shuai WangNeurIPS 2023 · 69 citations
- Do Users Benefit From Interpretable Vision? A User Study, Baseline, And DatasetLeon Sixt, Martin Schuessler, Oana-Iuliana Popescu, Philipp Weiß et al.ICLR 2022 · 21 citations
