Part-level Semantic-guided Contrastive Learning for Fine-grained Visual Classification
Zhijian Lin, Hong Han
Abstract
Fine-Grained Visual Classification (FGVC) aims to distinguish visually similar subcategories within a broad category, and poses significant challenges due to subtle inter-class differences, large intra-class variations, and data scarcity. Existing methods often struggle to effectively capture both part-level detail and spatial relational features, particularly across rigid and non-rigid object categories. To address these issues, we propose Part-level Semantic-guided Contrastive Learning (PSCL), a novel framework that integrates three key components. (1) The Part Localization Module (PLM) leverages clearCLIP to enable text-controllable region selection, achieving decoupled and semantically guided spatial feature extraction. (2) The Multi-scale Multi-part Branch Progressive Reasoning (MMBPR) module captures discriminative features across multiple parts and scales, while reducing inter-branch redundancy. (3) The Visual-Language Contrastive Learning based on Multi-grained Text Features (VLCL-MG) module introduces intermediate-granularity category concepts to improve feature alignment and inter-class separability. Extensive experiments on five publicly available FGVC datasets demonstrate the superior performance and generalization ability of PSCL, validating the effectiveness of its modular design and the synergy between vision and language. Code is available at: https://github.com/joker-lin9/PSCL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cedcfad3-8888-4d4d-970b-ab7d710945a8Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsSiyuan Li, Li Sun, Qingli LiAAAI 2023 · 355 citations
Related papers
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du et al.AAAI 2024 · 71 citations
- PaCL: Part-level Contrastive Learning for Fine-grained Few-shot Image ClassificationChuanming Wang, Huiyuan Fu, Huadong MaACM MM 2022 · 23 citations
- CLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised LearningXiaokun Li, Yaping Huang, Qingji GuanCVPR 2025
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
