Multi-Part Token Transformer with Dual Contrastive Learning for Fine-grained Image Classification
Chuanming Wang, Huiyuan Fu, Huadong Ma
摘要
Fine-grained image classification focuses on distinguishing objects from different similar subcategories, which requires the classification model to extract subtle yet discriminative descriptors. Recent Vision Transformer (ViT) has shown an enormous potential for this challenging task, but previous ViT-based methods have primarily focused on improving the relationship between image patches, neglecting the limited expressive capability caused by the single class token.To address this limitation, we propose to learn a Multi-part Token Transformer (MpT-Trans), which extends the class token to multiple tokens presenting various parts, enhancing the model's capability of extracting discriminative information. Specifically, our MpT-Trans model interpolates the vision transformer framework with two modules: (i) the Part-wise Shift Learning (PwSL) module is proposed to extend the single class token to a set of part tokens with differentiable shifts, enabling the model to extract informative representations from different perspectives; (ii) the Dual Contrastive Learning (DuCL) module is introduced to exploit the inter-class and inter-part relationships to regularize the learning of part tokens, enhancing their diversity and discrimination for accurate classification. Extensive experiments and ablation study demonstrate that the proposed MpT-Trans achieves state-of-the-art performance on various fine-grained image benchmark datasets, demonstrating the effectiveness of our proposed method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Multi-scale Activation, Selection, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird RecognitionZhicheng Zhang, Hao Tang, Jinhui TangAAAI 2025 · 被引用 6 次
- Disentangled Hypergraph-Guided Mamba Scanning for Fine-Grained Visual RecognitionZhongwei Xiong, Hao Wang, Xiaoyan Yu, Lingling Li 等AAAI 2026
- Part-level Semantic-guided Contrastive Learning for Fine-grained Visual ClassificationZhijian Lin, Hong HanICLR 2026
相关 Paper
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- SIM-Trans: Structure Information Modeling Transformer for Fine-grained Visual CategorizationHongbo Sun, Xiangteng He, Yuxin PengACM MM 2022 · 被引用 128 次
- PaCL: Part-level Contrastive Learning for Fine-grained Few-shot Image ClassificationChuanming Wang, Huiyuan Fu, Huadong MaACM MM 2022 · 被引用 23 次
- Token Contrast for Weakly-Supervised Semantic SegmentationLixiang Ru, Heliang Zheng, Yibing Zhan, Bo DuCVPR 2023
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd 等CVPR 2022 · 被引用 275 次
