ViTree: Single-Path Neural Tree for Step-Wise Interpretable Fine-Grained Visual Categorization
Danning Lao, Qi Liu, Jiazi Bu, Junchi Yan, Wei Shen
Abstract
As computer vision continues to advance and finds widespread applications across various domains, the need for interpretability in deep learning models becomes paramount. Existing methods often resort to post-hoc techniques or prototypes to explain the decision-making process, which can be indirect and lack intrinsic illustration. In this research, we introduce ViTree, a novel approach for fine-grained visual categorization that combines the popular vision transformer as a feature extraction backbone with neural decision trees. By traversing the tree paths, ViTree effectively selects patches from transformer-processed features to highlight informative local regions, thereby refining representations in a step-wise manner. Unlike previous tree-based models that rely on soft distributions or ensembles of paths, ViTree selects a single tree path, offering a clearer and simpler decision-making process. This patch and path selectivity enhances model interpretability of ViTree, enabling better insights into the model's inner workings. Remarkably, extensive experimentation validates that this streamlined approach surpasses various strong competitors and achieves state-of-the-art performance while maintaining exceptional interpretability which is proved by multi-perspective methods. Code can be found at https://github.com/SJTU-DeepVisionLab/ViTree .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 142d019c-ca9f-4df4-9056-a2bedfca9a74Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MetaFormer is Actually What You Need for VisionWeihao Yu, Mi Luo, Pan Zhou, Chenyang Si et al.CVPR 2022 · 1,114 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- SIM-Trans: Structure Information Modeling Transformer for Fine-grained Visual CategorizationHongbo Sun, Xiangteng He, Yuxin PengACM MM 2022 · 128 citations
Related papers
- ViT-NeT: Interpretable Vision Transformers with Neural Tree DecoderSangwon Kim, Jae-Yeal Nam, ByoungChul KoICML 2022 · 97 citations
- Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-AttentionSaebom Leem, Hyunseok SeoAAAI 2024 · 40 citations
- Attention Convolutional Binary Neural Tree for Fine-Grained Visual CategorizationRuyi Ji, Longyin Wen, Libo Zhang, Dawei Du et al.CVPR 2020
- CF-ViT: A General Coarse-to-Fine Method for Vision TransformerMengzhao Chen, Mingbao Lin, Ke Li, Yunhang Shen et al.AAAI 2023 · 105 citations
- Interpretable and Accurate Fine-grained Recognition via Region GroupingZixuan Huang, Yin LiCVPR 2020
