Multi-View Active Fine-Grained Visual Recognition
Ruoyi Du, Wenqing Yu, Heqing Wang, Ting-En Lin, Dongliang Chang, Zhanyu Ma
Abstract
Despite the remarkable progress of Fine-grained visual classification (FGVC) with years of history, it is still limited to recognizing 2D images. Recognizing objects in the physical world (i.e., 3D environment) poses a unique challenge – discriminative information is not only present in visible local regions but also in other unseen views. Therefore, in addition to finding the distinguishable part from the current view, efficient and accurate recognition requires inferring the critical perspective with minimal glances. E.g., a person might recognize a "Ford sedan" with a glance at its side and then know that looking at the front can help tell which model it is. In this paper, towards FGVC in the real physical world, we put forward the problem of multi-view active fine-grained visual recognition (MAFR) and complete this study in three steps: (i) a multi-view, fine-grained vehicle dataset is collected as the testbed, (ii) a pilot experiment is designed to validate the need and research value of MAFR, (iii) a policy-gradient-based framework along with a dynamic exiting strategy is proposed to achieve efficient recognition with active view selection. Our comprehensive experiments demonstrate that the proposed method outperforms previous multi-view recognition works and can extend existing state-of-the-art FGVC methods and advanced neural networks to become "FGVC experts" in the 3D environment. Our code is available at https://github.com/PRIS-CV/MAFR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e65e4edb-36ba-4c19-80d2-a2e304f824beCited by top-tier papers4
- Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual UnderstandingJunhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang et al.CVPR 2026 · 2 citations
- Switch-a-View: View Selection Learned from Unlabeled In-the-Wild VideosSagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Kristen GraumanICCV 2025 · 1 citation
- Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional VideosSagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan et al.CVPR 2025
- Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency PartitionLintong Zhang, Kang Yin, Seong-Whan LeeCVPR 2025
Builds on7
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- VV-Net: Voxel VAE Net With Group Convolutions for Point Cloud SegmentationHsien-Yu Meng, Lin Gao, Yu-Kun Lai, Dinesh ManochaICCV 2019 · 268 citations
- Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image ClassificationYulin Wang, Kangchen Lv, Rui Huang, Shiji Song et al.NeurIPS 2020 · 179 citations
- Learning Relationships for Multi-View 3D Object RecognitionZe Yang, Liwei WangICCV 2019 · 166 citations
- On-the-Fly Category DiscoveryRuoyi Du, Dongliang Chang, Kongming Liang, Timothy M. Hospedales et al.CVPR 2023
Related papers
- Learning Canonical 3D Object Representation for Fine-Grained RecognitionSunghun Joung, Seungryong Kim, Minsu Kim, Ig-Jae Kim et al.ICCV 2021 · 14 citations
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
- Beyond the Parts: Learning Multi-view Cross-part Correlation for Vehicle Re-identificationXinchen Liu, Wu Liu, Jinkai Zheng, Chenggang Yan et al.ACM MM 2020 · 97 citations
- Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesManuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne et al.ICCV 2019 · 235 citations
- An Erudite Fine-Grained Visual Classification ModelDongliang Chang, Yujun Tong, Ruoyi Du, Timothy M. Hospedales et al.CVPR 2023
