Self-Interpretable Model with Transformation Equivariant Interpretation
Yipei Wang, Xiaoqian Wang
Abstract
In this paper, we propose a self-interpretable model SITE with transformation-equivariant interpretations. We focus on the robustness and self-consistency of the interpretations of geometric transformations. Apart from the transformation equivariance, as a self-interpretable model, SITE has comparable expressive power as the benchmark black-box classifiers, while being able to present faithful and robust interpretations with high quality. It is worth noticing that although applied in most of the CNN visualization methods, the bilinear upsampling approximation is a rough approximation, which can only provide interpretations in the form of heatmaps (instead of pixel-wise). It remains an open question whether such interpretations can be direct to the input space (as shown in the MNIST experiments). Besides, we consider the translation and rotation transformations in our model. In future work, we will explore the robust interpretations under more complex transformations such as scaling and distortion. Moreover, we clarify that SITE is not limited to geometric transformation (that we used in the computer vision domain), and will explore SITEin other domains in future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86afdfc4-892d-4cf4-8703-477e2e8d99acCited by top-tier papers7
- ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational ModelSrishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin et al.NeurIPS 2022 · 55 citations
- Evaluating the Robustness of Interpretability Methods through Explanation Invariance and EquivarianceJonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 27 citations
- Path Choice Matters for Clear Attributions in Path MethodsBorui Zhang, Wenzhao Zheng, Jie Zhou, Jiwen LuICLR 2024 · 5 citations
- Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and RetrainingAaron Jiaxun Li, Robin Netzorg, Zhihan Cheng, Zhuoqin Zhang et al.ICML 2024 · 5 citations
- Benchmarking Deletion Metrics with the Principled ExplanationsYipei Wang, Xiaoqian WangICML 2024 · 4 citations
Builds on8
- Neural Additive Models: Interpretable Machine Learning with Neural NetsRishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang et al.NeurIPS 2021 · 663 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
- On Learning Sets of Symmetric ElementsHaggai Maron, Or Litany, Gal Chechik, Ethan FetayaICML 2020 · 148 citations
- Proper Network Interpretability Helps Adversarial Robustness in ClassificationAkhilan Boopathy, Sijia Liu, Gaoyuan Zhang, Cynthia Liu et al.ICML 2020 · 74 citations
Related papers
- A Disentangling Invertible Interpretation Network for Explaining Latent RepresentationsPatrick Esser, Robin Rombach, Björn OmmerCVPR 2020
- A Hypertoroidal Covering for Perfect Color EquivarianceYulong Yang, Zhikun Xu, Yaojun Li, Christine Allen-BlanchetteICML 2026
- Make it SING: Analyzing Semantic Invariants in ClassifiersHarel Yadid, Meir Yossef Levi, Roy Betser, Guy GilboaCVPR 2026
- REST: Performance Improvement of a Black Box Model via RL-Based Spatial TransformationJae-Myung Kim, Hyungjin Kim, Chanwoo Park, Jungwoo LeeAAAI 2020
- One step further: evaluating interpreters using metamorphic testingMing Fan, Jiali Wei, Wuxia Jin, Zhou Xu et al.ISSTA 2022 · 7 citations
