Learning Hierarchal Channel Attention for Fine-grained Visual Classification
Xiang Guan, Guoqing Wang, Xing Xu, Yi Bin
摘要
Learning delicate feature representation of object parts plays a critical role in fine-grained visual classification tasks. However, advanced deep convolutional neural networks trained for general visual classification tasks usually tend to focus on the coarse-grained information while ignoring the fine-grained one, which is of great significance for learning discriminative representation. In this work, we explore the great merit of multi-modal data in introducing semantic knowledge and sequential analysis techniques in learning hierarchical feature representation for generating discriminative fine-grained features. To this end, we propose a novel approach, termed Channel Cusum Attention ResNet (CCA-ResNet ), for multi-modal joint learning of fine-grained representation. Specifically, we use feature-level multi-modal alignment to connect image and text classification models for joint multi-modal training. Through joint training, image classification models trained with semantic level labels tend to focus on the most discriminative parts, which enhances the cognitive ability of the model. Then, we propose a Channel Cusum Attention (CCA ) mechanism to equip feature maps with hierarchical properties through unsupervised reconstruction of local and global features. The benefits brought by the CCA are in two folds: a) allowing fine-grained features from early layers to be preserved in the forward propagation of deep networks; b) leveraging the hierarchical properties to facilitate multi-modal feature alignment. We conduct extensive experiments to verify that our proposed model can achieve state-of-the-art performance on a series of fine-grained visual classification benchmarks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti 等NeurIPS 2022 · 被引用 302 次
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal SearchMeiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing 等AAAI 2024 · 被引用 24 次
- Category-specific Semantic Coherency Learning for Fine-grained Image RecognitionShijie Wang, Zhihui Wang, Haojie Li, Wanli OuyangACM MM 2020 · 被引用 23 次
- A Weakly Supervised Fine Label Classifier Enhanced by Coarse SupervisionFariborz Taherkhani, Hadi Kazemi, Ali Dabouei, Jeremy M. Dawson 等ICCV 2019 · 被引用 30 次
- Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental LearningXusheng Cao, Haori Lu, Linlan Huang, Fei Yang 等NeurIPS 2025 · 被引用 3 次
