Cross-X Learning for Fine-Grained Visual Categorization
Wei Luo, Xitong Yang, Xianjie Mo, Yuheng Lu, Larry Davis, Jun Li, Jian Yang, Ser-Nam Lim
Abstract
Recognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features are extracted for fine-grained classification. However, these methods typically treat the part-specific features of each image in isolation while neglecting their relationships between different images. In this paper, we propose Cross-X learning, a simple yet effective approach that exploits the relationships between different images and between different network layers for robust multi-scale feature learning. Our approach involves two novel components: (i) a cross-category cross-semantic regularizer that guides the extracted features to represent semantic parts and, (ii) a cross-layer regularizer that improves the robustness of multi-scale features by matching the prediction distribution across multiple layers. Our approach can be easily trained end-to-end and is scalable to large datasets like NABirds. We empirically analyze the contributions of different components of our approach and demonstrate its robustness, effectiveness and state-of-the-art performance on five benchmark datasets. Code is available at https://github.com/cswluo/CrossX.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23607647-45a8-4a07-96a2-e81766051e95Cited by top-tier papers21
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu et al.CVPR 2022 · 251 citations
- Context-aware Attentional Pooling (CAP) for Fine-grained Visual ClassificationArdhendu Behera, Zachary Wharton, Pradeep R. P. G. Hewage, Asish BeraAAAI 2021 · 142 citations
- RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionYunqing Hu, Xuan Jin, Yin Zhang, Haiwen Hong et al.ACM MM 2021 · 142 citations
- ViT-NeT: Interpretable Vision Transformers with Neural Tree DecoderSangwon Kim, Jae-Yeal Nam, ByoungChul KoICML 2022 · 97 citations
- Fine-Grained Object Classification via Self-Supervised Pose AlignmentXuhui Yang, Yaowei Wang, Ke Chen, Yong Xu et al.CVPR 2022 · 91 citations
Related papers
- Unsupervised Part Discovery from Contrastive ReconstructionSubhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2021 · 74 citations
- Weakly-Supervised Semantic Segmentation via Sub-Category ExplorationYu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu et al.CVPR 2020
- Learning Attribute and Class-Specific Representation Duet for Fine-Grained Fashion AnalysisYang Jiao, Yan Gao, Jingjing Meng, Jin Shang et al.CVPR 2023
- Global Meets Local: Effective Multi-Label Image Classification via Category-Aware Weak SupervisionJiawei Zhan, Jun Liu, Wei Tang, Guannan Jiang et al.ACM MM 2022 · 6 citations
- Informative and Consistent Correspondence Mining for Cross-Domain Weakly Supervised Object DetectionLuwei Hou, Yu Zhang, Kui Fu, Jia LiCVPR 2021
