An Erudite Fine-Grained Visual Classification Model
Dongliang Chang, Yujun Tong, Ruoyi Du, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma
Abstract
Current fine-grained visual classification (FGVC) models are isolated. In practice, we first need to identify the coarse-grained label of an object, then select the corresponding FGVC model for recognition. This hinders the application of FGVC algorithms in real-life scenarios. In this paper, we propose an erudite FGVC model jointly trained by several different datasets 1 , which can efficiently and accurately predict an object's fine-grained label across the combined label space. We found through a pilot study that positive and negative transfers co-occur when different datasets are mixed for training, i.e., the knowledge from other datasets is not always useful. Therefore, we first propose a feature disentanglement module and a feature re-fusion module to reduce negative transfer and boost positive transfer between different datasets. In detail, we reduce negative transfer by decoupling the deep features through many dataset-specific feature extractors. Subsequently, these are channel-wise re-fused to facilitate positive transfer. Finally, we propose a meta-learning based dataset-agnostic spatial attention layer to take full advantage of the multi-dataset training data, given that localisation is dataset-agnostic between different datasets. Experimental results across 11 different mixed-datasets built on four different FGVC datasets demonstrate the effectiveness of the proposed method. Furthermore, the proposed method can be easily combined with existing FGVC methods to obtain state-of-the-art results. Our code is available at https://github.com/PRIS-CV/An-Erudite- FGVC-Model. * indicates the corresponding author. 1 In this paper, different datasets mean different fine-grained visual classification datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Multi-View Active Fine-Grained Visual RecognitionRuoyi Du, Wenqing Yu, Heqing Wang, Ting-En Lin et al.ICCV 2023 · 14 citations
- Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous LearnerHanwen Zhong, Jiaxin Chen, Yutong Zhang, Di Huang et al.NeurIPS 2024 · 9 citations
- Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual UnderstandingJunhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang et al.CVPR 2026 · 2 citations
Builds on9
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Learning Attentive Pairwise Interaction for Fine-Grained ClassificationPeiqin Zhuang, Yali Wang, Yu QiaoAAAI 2020 · 392 citations
- Fine-Grained Recognition: Accounting for Subtle Differences between Similar ClassesGuolei Sun, Hisham Cholakkal, Salman H. Khan, Fahad Shahbaz Khan et al.AAAI 2020 · 138 citations
- Simple Multi-dataset DetectionXingyi Zhou, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 85 citations
Related papers
- Your "Flamingo" is My "Bird": Fine-Grained, or NotDongliang Chang, Kaiyue Pang, Yixiao Zheng, Zhanyu Ma et al.CVPR 2021
- Data-free Knowledge Distillation for Fine-grained Visual CategorizationRenrong Shao, Wei Zhang, Jianhua Yin, Jun WangICCV 2023 · 7 citations
- mDALU: Multi-Source Domain Adaptation and Label Unification with Partial DatasetsRui Gong, Dengxin Dai, Yuhua Chen, Wen Li et al.ICCV 2021 · 27 citations
- Filtration and Distillation: Enhancing Region Attention for Fine-Grained Visual CategorizationChuanbin Liu, Hongtao Xie, Zheng-Jun Zha, Lingfeng Ma et al.AAAI 2020 · 179 citations
- SIM-Trans: Structure Information Modeling Transformer for Fine-grained Visual CategorizationHongbo Sun, Xiangteng He, Yuxin PengACM MM 2022 · 128 citations
