Self-Interpretable Model with Transformation Equivariant Interpretation
Yipei Wang, Xiaoqian Wang
摘要
In this paper, we propose a self-interpretable model SITE with transformation-equivariant interpretations. We focus on the robustness and self-consistency of the interpretations of geometric transformations. Apart from the transformation equivariance, as a self-interpretable model, SITE has comparable expressive power as the benchmark black-box classifiers, while being able to present faithful and robust interpretations with high quality. It is worth noticing that although applied in most of the CNN visualization methods, the bilinear upsampling approximation is a rough approximation, which can only provide interpretations in the form of heatmaps (instead of pixel-wise). It remains an open question whether such interpretations can be direct to the input space (as shown in the MNIST experiments). Besides, we consider the translation and rotation transformations in our model. In future work, we will explore the robust interpretations under more complex transformations such as scaling and distortion. Moreover, we clarify that SITE is not limited to geometric transformation (that we used in the computer vision domain), and will explore SITEin other domains in future work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational ModelSrishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin 等NeurIPS 2022 · 被引用 55 次
- Evaluating the Robustness of Interpretability Methods through Explanation Invariance and EquivarianceJonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 被引用 27 次
- Path Choice Matters for Clear Attributions in Path MethodsBorui Zhang, Wenzhao Zheng, Jie Zhou, Jiwen LuICLR 2024 · 被引用 5 次
- Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and RetrainingAaron Jiaxun Li, Robin Netzorg, Zhihan Cheng, Zhuoqin Zhang 等ICML 2024 · 被引用 5 次
- Benchmarking Deletion Metrics with the Principled ExplanationsYipei Wang, Xiaoqian WangICML 2024 · 被引用 4 次
它引用的顶会 Paper8
- Neural Additive Models: Interpretable Machine Learning with Neural NetsRishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang 等NeurIPS 2021 · 被引用 663 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 被引用 249 次
- On Learning Sets of Symmetric ElementsHaggai Maron, Or Litany, Gal Chechik, Ethan FetayaICML 2020 · 被引用 148 次
- Proper Network Interpretability Helps Adversarial Robustness in ClassificationAkhilan Boopathy, Sijia Liu, Gaoyuan Zhang, Cynthia Liu 等ICML 2020 · 被引用 74 次
相关 Paper
- A Disentangling Invertible Interpretation Network for Explaining Latent RepresentationsPatrick Esser, Robin Rombach, Björn OmmerCVPR 2020
- A Hypertoroidal Covering for Perfect Color EquivarianceYulong Yang, Zhikun Xu, Yaojun Li, Christine Allen-BlanchetteICML 2026
- Make it SING: Analyzing Semantic Invariants in ClassifiersHarel Yadid, Meir Yossef Levi, Roy Betser, Guy GilboaCVPR 2026
- REST: Performance Improvement of a Black Box Model via RL-Based Spatial TransformationJae-Myung Kim, Hyungjin Kim, Chanwoo Park, Jungwoo LeeAAAI 2020
- One step further: evaluating interpreters using metamorphic testingMing Fan, Jiali Wei, Wuxia Jin, Zhou Xu 等ISSTA 2022 · 被引用 7 次
