When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object Detection
Hongliang Zhou, Yongxiang Liu, Canyu Mo, Weijie Li, Bowen Peng, Li Liu
Abstract
Few-shot object detection aims to detect novel classes with limited samples. Recent methods have leveraged rich semantic representations of pretrained vision transformer (ViT) to overcome limitations of model fine-tuning, thereby improving performance on novel classes. However, existing pretrained ViT schemes only perform transformer encoding in feature dimension, ignoring exploration of pixel-wise differences in low-level features and multiscale variations. The current challenges lie in: (i) extracted features suffer from blurred boundary features and smooth transition from center to boundary, leading to insufficient distinction between objects and backgrounds, and (ii) how to balance extraction of local details and global contour features under multiscale scenarios. So Pixel Difference Vision Transformer (PiDiViT) is proposed. Innovations include: (i) difference convolution fusion module (DCFM), which enhances feature differences from object centers to boundaries and effectively preserves global information by fusing pixel-wise central difference features with original features through an attention mechanism, and (ii) multiscale feature fusion module (MFFM), which adaptively fuses features extracted by five different scale convolutional kernels using a scale attention mechanism to generate attention weights, achieving an optimal balance between local detail and global semantic information extraction. PiDiViT achieves SOTA on the COCO benchmark: surpassing few-shot detection SOTA by 2.7 nAP50 (10-shot) and 4.0 nAP50 (30-shot) for novel classes, exceeding one-shot detection SOTA by 4.4 nAP50 and open-vocabulary detection SOTA by 3.7 nAP50. The code is available at https://github.com/Seaz9/PiDiViT.
- CSPS Building upon DE-ViT [49], we replace the dot product of DE-ViT [49] with cosine similarity [51] to calculate projection between category prototypes and ViT features to obtain input features. It mitigates overfitting caused by sample imbalance between novel and base classes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell et al.ICML 2020 · 723 citations
- Meta R-CNN: Towards General Solver for Instance-Level Low-Shot LearningXiaopeng Yan, Ziliang Chen, Anni Xu, Xiaoxi Wang et al.ICCV 2019 · 590 citations
Related papers
- FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingAdrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 61 citations
- PS-TTL: Prototype-based Soft-labels and Test-Time Learning for Few-shot Object DetectionYingjie Gao, Yanan Zhang, Ziyue Huang, Nanqing Liu et al.ACM MM 2024 · 13 citations
- Exploring Effective Knowledge Transfer for Few-shot Object DetectionZhiyuan Zhao, Qingjie Liu, Yunhong WangACM MM 2022 · 16 citations
- Few-Shot Object Detection with Fully Cross-TransformerGuangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen et al.CVPR 2022 · 183 citations
- Few-Shot Object Detection via Association and DIscriminationYuhang Cao, Jiaqi Wang, Ying Jin, Tong Wu et al.NeurIPS 2021 · 110 citations
