Learning Hierarchal Channel Attention for Fine-grained Visual Classification
Xiang Guan, Guoqing Wang, Xing Xu, Yi Bin
Abstract
Learning delicate feature representation of object parts plays a critical role in fine-grained visual classification tasks. However, advanced deep convolutional neural networks trained for general visual classification tasks usually tend to focus on the coarse-grained information while ignoring the fine-grained one, which is of great significance for learning discriminative representation. In this work, we explore the great merit of multi-modal data in introducing semantic knowledge and sequential analysis techniques in learning hierarchical feature representation for generating discriminative fine-grained features. To this end, we propose a novel approach, termed Channel Cusum Attention ResNet (CCA-ResNet ), for multi-modal joint learning of fine-grained representation. Specifically, we use feature-level multi-modal alignment to connect image and text classification models for joint multi-modal training. Through joint training, image classification models trained with semantic level labels tend to focus on the most discriminative parts, which enhances the cognitive ability of the model. Then, we propose a Channel Cusum Attention (CCA ) mechanism to equip feature maps with hierarchical properties through unsupervised reconstruction of local and global features. The benefits brought by the CCA are in two folds: a) allowing fine-grained features from early layers to be preserved in the forward propagation of deep networks; b) leveraging the hierarchical properties to facilitate multi-modal feature alignment. We conduct extensive experiments to verify that our proposed model can achieve state-of-the-art performance on a series of fine-grained visual classification benchmarks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get aadbe1b0-accc-4b47-82e8-f4e04fae3341Related papers
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti et al.NeurIPS 2022 · 302 citations
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal SearchMeiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing et al.AAAI 2024 · 24 citations
- Category-specific Semantic Coherency Learning for Fine-grained Image RecognitionShijie Wang, Zhihui Wang, Haojie Li, Wanli OuyangACM MM 2020 · 23 citations
- A Weakly Supervised Fine Label Classifier Enhanced by Coarse SupervisionFariborz Taherkhani, Hadi Kazemi, Ali Dabouei, Jeremy M. Dawson et al.ICCV 2019 · 30 citations
- Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental LearningXusheng Cao, Haori Lu, Linlan Huang, Fei Yang et al.NeurIPS 2025 · 3 citations
