ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder
Sangwon Kim, Jae-Yeal Nam, ByoungChul Ko
摘要
Vision transformers (ViTs), which have demonstrated a state-of-the-art performance in image classification, can also visualize global interpretations through attention-based contributions. However, the complexity of the model makes it difficult to interpret the decision-making process, and the ambiguity of the attention maps can cause incorrect correlations between image patches. In this study, we propose a new ViT neural tree decoder (ViT-NeT). A ViT acts as a backbone, and to solve its limitations, the output contextual image patches are applied to the proposed NeT. The NeT aims to accurately classify finegrained objects with similar inter-class correlations and different intra-class correlations. In addition, it describes the decision-making process through a tree structure and prototype and enables a visual interpretation of the results. The proposed ViT-NeT is designed to not only improve the classification performance but also provide a human-friendly interpretation, which is effective in resolving the trade-off between performance and interpretability. We compared the performance of ViT-NeT with other state-of-art methods using widely used fine-grained visual categorization benchmark datasets and experimentally proved that the proposed method is superior in terms of the classification performance and interpretability. The code and models are publicly available at https://github.com/ jumpsnack/ViT-NeT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Learning Support and Trivial Prototypes for Interpretable Image ClassificationChong Wang, Yuyuan Liu, Yuanhong Chen, Fengbei Liu 等ICCV 2023 · 被引用 50 次
- Interpretable Image Classification with Adaptive Prototype-based Vision TransformersChiyu Ma, Jon Donnelly, Wenjun Liu, Soroush Vosoughi 等NeurIPS 2024 · 被引用 48 次
- A Simple Interpretable Transformer for Fine-Grained Image Classification and AnalysisDipanjyoti Paul, Arpita Chowdhury, Xinqi Xiong, Feng-Ju Chang 等ICLR 2024 · 被引用 27 次
- Learning Time in Static ClassifiersXi Ding, Lei Wang, Piotr Koniusz, Yongsheng GaoAAAI 2026 · 被引用 2 次
- ViTree: Single-Path Neural Tree for Step-Wise Interpretable Fine-Grained Visual CategorizationDanning Lao, Qi Liu, Jiazi Bu, Junchi Yan 等AAAI 2024 · 被引用 1 次
它引用的顶会 Paper11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- Learning Attentive Pairwise Interaction for Fine-Grained ClassificationPeiqin Zhuang, Yali Wang, Yu QiaoAAAI 2020 · 被引用 392 次
相关 Paper
- Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-AttentionSaebom Leem, Hyunseok SeoAAAI 2024 · 被引用 40 次
- Attention Convolutional Binary Neural Tree for Fine-Grained Visual CategorizationRuyi Ji, Longyin Wen, Libo Zhang, Dawei Du 等CVPR 2020
- MG-ViT: A Multi-Granularity Method for Compact and Efficient Vision TransformersYu Zhang, Yepeng Liu, Duoqian Miao, Qi Zhang 等NeurIPS 2023 · 被引用 23 次
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He 等ICCV 2021 · 被引用 154 次
- SIM-Trans: Structure Information Modeling Transformer for Fine-grained Visual CategorizationHongbo Sun, Xiangteng He, Yuxin PengACM MM 2022 · 被引用 128 次
