Adaptive Image Transformer for One-Shot Object Detection
Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu
Abstract
One-shot object detection tackles a challenging task that aims at identifying within a target image all object instances of the same class, implied by a query image patch. The main difficulty lies in the situation that the class label of the query patch and its respective examples are not available in the training data. Our main idea leverages the concept of language translation to boost metric-learning-based detection methods. Specifically, we emulate the language translation process to adaptively translate the feature of each object proposal to better correlate the given query feature for discriminating the class-similarity among the proposal-query pairs. To this end, we propose the Adaptive Image Transformer (AIT) module that deploys an attentionbased encoder-decoder architecture to simultaneously explore intra-coder and inter-coder (i.e., each proposal-query pair) attention. The adaptive nature of our design turns out to be flexible and effective in addressing the one-shot learning scenario. With the informative attention cues, the proposed model excels in predicting the class-similarity between the target image proposals and the query image patch. Though conceptually simple, our model significantly outperforms a state-of-the-art technique, improving the unseenclass object classification from 63.8 mAP and 22.0 AP50 to 72.2 mAP and 24.3 AP50 on the PASCAL-VOC and MS-COCO benchmark datasets, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Few-Shot Object Detection with Fully Cross-TransformerGuangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen et al.CVPR 2022 · 183 citations
- Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular VideoWen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark LiaoCVPR 2022 · 117 citations
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 69 citations
- Semantic-aligned Fusion Transformer for One-shot Object DetectionYizhou Zhao, Xun Guo, Yan LuCVPR 2022 · 29 citations
- Balanced and Hierarchical Relation Learning for One-shot Object DetectionHanqing Yang, Sijia Cai, Hualian Sheng, Bing Deng et al.CVPR 2022 · 28 citations
Builds on5
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- Dynamic Anchor Feature Selection for Single-Shot Object DetectionShuai Li, Lingxiao Yang, Jianqiang Huang, Xian-Sheng Hua et al.ICCV 2019 · 52 citations
- Few-Shot Object Detection With Attention-RPN and Multi-Relation DetectorQi Fan, Wei Zhuo, Chi-Keung Tang, Yu-Wing TaiCVPR 2020
- Rethinking Classification and Localization for Object DetectionYue Wu, Yinpeng Chen, Lu Yuan, Zicheng Liu et al.CVPR 2020
- Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample SelectionShifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei et al.CVPR 2020
Related papers
- QDETRv: Query-Guided DETR for One-Shot Object Localization in VideosYogesh Kumar, Saswat Mallick, Anand Mishra, Sowmya Rasipuram et al.AAAI 2024 · 4 citations
- UP-DETR: Unsupervised Pre-Training for Object Detection With TransformersZhigang Dai, Bolun Cai, Yugeng Lin, Junying ChenCVPR 2021
- Reformulating HOI Detection As Adaptive Set PredictionMingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen et al.CVPR 2021
- CAD: Co-Adapting Discriminative Features for Improved Few-Shot ClassificationPhilip Chikontwe, Soopil Kim, Sang Hyun ParkCVPR 2022 · 46 citations
- Dynamic Transformer for Few-shot Instance SegmentationHaochen Wang, Jie Liu, Yongtuo Liu, Subhransu Maji et al.ACM MM 2022 · 10 citations
