When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object Detection
Hongliang Zhou, Yongxiang Liu, Canyu Mo, Weijie Li, Bowen Peng, Li Liu
摘要
Few-shot object detection aims to detect novel classes with limited samples. Recent methods have leveraged rich semantic representations of pretrained vision transformer (ViT) to overcome limitations of model fine-tuning, thereby improving performance on novel classes. However, existing pretrained ViT schemes only perform transformer encoding in feature dimension, ignoring exploration of pixel-wise differences in low-level features and multiscale variations. The current challenges lie in: (i) extracted features suffer from blurred boundary features and smooth transition from center to boundary, leading to insufficient distinction between objects and backgrounds, and (ii) how to balance extraction of local details and global contour features under multiscale scenarios. So Pixel Difference Vision Transformer (PiDiViT) is proposed. Innovations include: (i) difference convolution fusion module (DCFM), which enhances feature differences from object centers to boundaries and effectively preserves global information by fusing pixel-wise central difference features with original features through an attention mechanism, and (ii) multiscale feature fusion module (MFFM), which adaptively fuses features extracted by five different scale convolutional kernels using a scale attention mechanism to generate attention weights, achieving an optimal balance between local detail and global semantic information extraction. PiDiViT achieves SOTA on the COCO benchmark: surpassing few-shot detection SOTA by 2.7 nAP50 (10-shot) and 4.0 nAP50 (30-shot) for novel classes, exceeding one-shot detection SOTA by 4.4 nAP50 and open-vocabulary detection SOTA by 3.7 nAP50. The code is available at https://github.com/Seaz9/PiDiViT.
- CSPS Building upon DE-ViT [49], we replace the dot product of DE-ViT [49] with cosine similarity [51] to calculate projection between category prototypes and ViT features to obtain input features. It mitigates overfitting caused by sample imbalance between novel and base classes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu 等ICCV 2019 · 被引用 835 次
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell 等ICML 2020 · 被引用 723 次
- Meta R-CNN: Towards General Solver for Instance-Level Low-Shot LearningXiaopeng Yan, Ziliang Chen, Anni Xu, Xiaoxi Wang 等ICCV 2019 · 被引用 590 次
相关 Paper
- FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingAdrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 被引用 61 次
- PS-TTL: Prototype-based Soft-labels and Test-Time Learning for Few-shot Object DetectionYingjie Gao, Yanan Zhang, Ziyue Huang, Nanqing Liu 等ACM MM 2024 · 被引用 13 次
- Exploring Effective Knowledge Transfer for Few-shot Object DetectionZhiyuan Zhao, Qingjie Liu, Yunhong WangACM MM 2022 · 被引用 16 次
- Few-Shot Object Detection with Fully Cross-TransformerGuangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen 等CVPR 2022 · 被引用 183 次
- Few-Shot Object Detection via Association and DIscriminationYuhang Cao, Jiaqi Wang, Ying Jin, Tong Wu 等NeurIPS 2021 · 被引用 110 次
