Cross-X Learning for Fine-Grained Visual Categorization
Wei Luo, Xitong Yang, Xianjie Mo, Yuheng Lu, Larry Davis, Jun Li, Jian Yang, Ser-Nam Lim
摘要
Recognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features are extracted for fine-grained classification. However, these methods typically treat the part-specific features of each image in isolation while neglecting their relationships between different images. In this paper, we propose Cross-X learning, a simple yet effective approach that exploits the relationships between different images and between different network layers for robust multi-scale feature learning. Our approach involves two novel components: (i) a cross-category cross-semantic regularizer that guides the extracted features to represent semantic parts and, (ii) a cross-layer regularizer that improves the robustness of multi-scale features by matching the prediction distribution across multiple layers. Our approach can be easily trained end-to-end and is scalable to large datasets like NABirds. We empirically analyze the contributions of different components of our approach and demonstrate its robustness, effectiveness and state-of-the-art performance on five benchmark datasets. Code is available at https://github.com/cswluo/CrossX.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu 等CVPR 2022 · 被引用 251 次
- Context-aware Attentional Pooling (CAP) for Fine-grained Visual ClassificationArdhendu Behera, Zachary Wharton, Pradeep R. P. G. Hewage, Asish BeraAAAI 2021 · 被引用 142 次
- RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionYunqing Hu, Xuan Jin, Yin Zhang, Haiwen Hong 等ACM MM 2021 · 被引用 142 次
- ViT-NeT: Interpretable Vision Transformers with Neural Tree DecoderSangwon Kim, Jae-Yeal Nam, ByoungChul KoICML 2022 · 被引用 97 次
- Fine-Grained Object Classification via Self-Supervised Pose AlignmentXuhui Yang, Yaowei Wang, Ke Chen, Yong Xu 等CVPR 2022 · 被引用 91 次
相关 Paper
- Unsupervised Part Discovery from Contrastive ReconstructionSubhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2021 · 被引用 74 次
- Weakly-Supervised Semantic Segmentation via Sub-Category ExplorationYu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu 等CVPR 2020
- Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion AnalysisYang Jiao, Yan Gao, Jingjing Meng, Jin Shang 等CVPR 2023
- Global Meets Local: Effective Multi-Label Image Classification via Category-Aware Weak SupervisionJiawei Zhan, Jun Liu, Wei Tang, Guannan Jiang 等ACM MM 2022 · 被引用 6 次
- Informative and Consistent Correspondence Mining for Cross-Domain Weakly Supervised Object DetectionLuwei Hou, Yu Zhang, Kui Fu, Jia LiCVPR 2021
