Learning Attentive Pairwise Interaction for Fine-Grained Classification
Peiqin Zhuang, Yali Wang, Yu Qiao
Abstract
Fine-grained classification is a challenging problem, due to subtle differences among highly-confused categories. Most approaches address this difficulty by learning discriminative representation of individual input image. On the other hand, humans can effectively identify contrastive clues by comparing image pairs. Inspired by this fact, this paper proposes a simple but effective Attentive Pairwise Interaction Network (API-Net), which can progressively recognize a pair of fine-grained images by interaction. Specifically, API-Net first learns a mutual feature vector to capture semantic differences in the input pair. It then compares this mutual vector with individual vectors to generate gates for each input image. These distinct gate vectors inherit mutual context on semantic differences, which allow API-Net to attentively capture contrastive clues by pairwise interaction between two images. Additionally, we train API-Net in an end-to-end manner with a score ranking regularization, which can further generalize API-Net by taking feature priorities into account. We conduct extensive experiments on five popular benchmarks in fine-grained classification. API-Net outperforms the recent SOTA methods, i.e., CUB-200-2011 (90.0%), Aircraft (93.9%), Stanford Cars (95.3%), Stanford Dogs (90.3%), and NABirds (88.1%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 330 citations
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu et al.CVPR 2022 · 251 citations
- RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionYunqing Hu, Xuan Jin, Yin Zhang, Haiwen Hong et al.ACM MM 2021 · 142 citations
- ViT-NeT: Interpretable Vision Transformers with Neural Tree DecoderSangwon Kim, Jae-Yeal Nam, ByoungChul KoICML 2022 · 97 citations
Related papers
- Channel Interaction Networks for Fine-Grained Image CategorizationYu Gao, Xintong Han, Xun Wang, Weilin Huang et al.AAAI 2020 · 177 citations
- Class Guided Channel Weighting Network for Fine-Grained Semantic SegmentationXiang Zhang, Wanqing Zhao, Hangzai Luo, Jinye Peng et al.AAAI 2022 · 3 citations
- Dual Attention Networks for Few-Shot Fine-Grained RecognitionShu-Lin Xu, Faen Zhang, Xiu-Shen Wei, Jianhua WangAAAI 2022 · 43 citations
- Dynamic Position-aware Network for Fine-grained Image RecognitionShijie Wang, Haojie Li, Zhihui Wang, Wanli OuyangAAAI 2021 · 36 citations
- Hierarchical Reasoning Network with Contrastive Learning for Few-Shot Human-Object Interaction RecognitionJiale Yu, Baopeng Zhang, Qirui Li, Haoyang Chen et al.ACM MM 2023 · 3 citations
