Context-aware Attentional Pooling (CAP) for Fine-grained Visual Classification
Ardhendu Behera, Zachary Wharton, Pradeep R. P. G. Hewage, Asish Bera
摘要
Deep convolutional neural networks (CNNs) have shown a strong ability in mining discriminative object pose and parts information for image recognition. For fine-grained recognition, context-aware rich feature representation of object/scene plays a key role since it exhibits a significant variance in the same subcategory and subtle variance among different subcategories. Finding the subtle variance that fully characterizes the object/scene is not straightforward. To address this, we propose a novel context-aware attentional pooling (CAP) that effectively captures subtle changes via sub-pixel gradients, and learns to attend informative integral regions and their importance in discriminating different subcategories without requiring the bounding-box and/or distinguishable part annotations. We also introduce a novel feature encoding by considering the intrinsic consistency between the informativeness of the integral regions and their spatial structures to capture the semantic correlation among them. Our approach is simple yet extremely effective and can be easily applied on top of a standard classification backbone network. We evaluate our approach using six state-of-the-art (SotA) backbone networks and eight benchmark datasets. Our method significantly outperforms the SotA approaches on six datasets and is very competitive with the remaining two.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Dynamic MLP for Fine-Grained Image Classification by Leveraging Geographical and Temporal InformationLingfeng Yang, Xiang Li, Renjie Song, Borui Zhao 等CVPR 2022 · 被引用 44 次
- Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive EvaluationHong-Tao Yu, Yuxin Peng, Serge J. Belongie, Xiu-Shen WeiICLR 2026 · 被引用 21 次
- FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object DetectionKe Li, Di Wang, Zhangyuan Hu, Shaofeng Li 等AAAI 2025 · 被引用 19 次
- Transitivity Recovering Decompositions: Interpretable and Robust Fine-Grained RelationshipsAbhra Chaudhuri, Massimiliano Mancini, Zeynep Akata, Anjan DuttaNeurIPS 2023 · 被引用 5 次
- Relational Proxies: Emergent Relationships as Fine-Grained DiscriminatorsAbhra Chaudhuri, Massimiliano Mancini, Zeynep Akata, Anjan DuttaNeurIPS 2022 · 被引用 5 次
它引用的顶会 Paper8
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani 等ICCV 2019 · 被引用 1,149 次
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
- Cross-X Learning for Fine-Grained Visual CategorizationWei Luo, Xitong Yang, Xianjie Mo, Yuheng Lu 等ICCV 2019 · 被引用 234 次
相关 Paper
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- Class Guided Channel Weighting Network for Fine-Grained Semantic SegmentationXiang Zhang, Wanqing Zhao, Hangzai Luo, Jinye Peng 等AAAI 2022 · 被引用 3 次
- Background-Aware Pooling and Noise-Aware Loss for Weakly-Supervised Semantic SegmentationYoungmin Oh, Beomjun Kim, Bumsub HamCVPR 2021
- Information Entropy Based Feature Pooling for Convolutional Neural NetworksWeitao Wan, Jiansheng Chen, Tianpeng Li, Yiqing Huang 等ICCV 2019 · 被引用 33 次
- Selective Sparse Sampling for Fine-Grained Image RecognitionYao Ding, Yanzhao Zhou, Yi Zhu, Qixiang Ye 等ICCV 2019 · 被引用 227 次
