Multi-View Active Fine-Grained Visual Recognition
Ruoyi Du, Wenqing Yu, Heqing Wang, Ting-En Lin, Dongliang Chang, Zhanyu Ma
摘要
Despite the remarkable progress of Fine-grained visual classification (FGVC) with years of history, it is still limited to recognizing 2D images. Recognizing objects in the physical world (i.e., 3D environment) poses a unique challenge – discriminative information is not only present in visible local regions but also in other unseen views. Therefore, in addition to finding the distinguishable part from the current view, efficient and accurate recognition requires inferring the critical perspective with minimal glances. E.g., a person might recognize a "Ford sedan" with a glance at its side and then know that looking at the front can help tell which model it is. In this paper, towards FGVC in the real physical world, we put forward the problem of multi-view active fine-grained visual recognition (MAFR) and complete this study in three steps: (i) a multi-view, fine-grained vehicle dataset is collected as the testbed, (ii) a pilot experiment is designed to validate the need and research value of MAFR, (iii) a policy-gradient-based framework along with a dynamic exiting strategy is proposed to achieve efficient recognition with active view selection. Our comprehensive experiments demonstrate that the proposed method outperforms previous multi-view recognition works and can extend existing state-of-the-art FGVC methods and advanced neural networks to become "FGVC experts" in the 3D environment. Our code is available at https://github.com/PRIS-CV/MAFR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual UnderstandingJunhan Chen, Zilu Zhou, Yujun Tong, Dongliang Chang 等CVPR 2026 · 被引用 2 次
- Switch-a-View: View Selection Learned from Unlabeled In-the-Wild VideosSagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Kristen GraumanICCV 2025 · 被引用 1 次
- Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional VideosSagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan 等CVPR 2025
- Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency PartitionLintong Zhang, Kang Yin, Seong-Whan LeeCVPR 2025
它引用的顶会 Paper7
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- VV-Net: Voxel VAE Net With Group Convolutions for Point Cloud SegmentationHsien-Yu Meng, Lin Gao, Yu-Kun Lai, Dinesh ManochaICCV 2019 · 被引用 268 次
- Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image ClassificationYulin Wang, Kangchen Lv, Rui Huang, Shiji Song 等NeurIPS 2020 · 被引用 179 次
- Learning Relationships for Multi-View 3D Object RecognitionZe Yang, Liwei WangICCV 2019 · 被引用 166 次
- On-the-Fly Category DiscoveryRuoyi Du, Dongliang Chang, Kongming Liang, Timothy M. Hospedales 等CVPR 2023
相关 Paper
- Learning Canonical 3D Object Representation for Fine-Grained RecognitionSunghun Joung, Seungryong Kim, Minsu Kim, Ig-Jae Kim 等ICCV 2021 · 被引用 14 次
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 被引用 37 次
- Beyond the Parts: Learning Multi-view Cross-part Correlation for Vehicle Re-identificationXinchen Liu, Wu Liu, Jinkai Zheng, Chenggang Yan 等ACM MM 2020 · 被引用 97 次
- Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesManuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne 等ICCV 2019 · 被引用 235 次
- An Erudite Fine-Grained Visual Classification ModelDongliang Chang, Yujun Tong, Ruoyi Du, Timothy M. Hospedales 等CVPR 2023
