Lune

ICCV2025Top-tier venue

When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object Detection

Hongliang Zhou, Yongxiang Liu, Canyu Mo, Weijie Li, Bowen Peng, Li Liu

2025Year
3Citations

Abstract

Few-shot object detection aims to detect novel classes with limited samples. Recent methods have leveraged rich semantic representations of pretrained vision transformer (ViT) to overcome limitations of model fine-tuning, thereby improving performance on novel classes. However, existing pretrained ViT schemes only perform transformer encoding in feature dimension, ignoring exploration of pixel-wise differences in low-level features and multiscale variations. The current challenges lie in: (i) extracted features suffer from blurred boundary features and smooth transition from center to boundary, leading to insufficient distinction between objects and backgrounds, and (ii) how to balance extraction of local details and global contour features under multiscale scenarios. So Pixel Difference Vision Transformer (PiDiViT) is proposed. Innovations include: (i) difference convolution fusion module (DCFM), which enhances feature differences from object centers to boundaries and effectively preserves global information by fusing pixel-wise central difference features with original features through an attention mechanism, and (ii) multiscale feature fusion module (MFFM), which adaptively fuses features extracted by five different scale convolutional kernels using a scale attention mechanism to generate attention weights, achieving an optimal balance between local detail and global semantic information extraction. PiDiViT achieves SOTA on the COCO benchmark: surpassing few-shot detection SOTA by 2.7 nAP50 (10-shot) and 4.0 nAP50 (30-shot) for novel classes, exceeding one-shot detection SOTA by 4.4 nAP50 and open-vocabulary detection SOTA by 3.7 nAP50. The code is available at https://github.com/Seaz9/PiDiViT.

  • CSPS Building upon DE-ViT [49], we replace the dot product of DE-ViT [49] with cosine similarity [51] to calculate projection between category prototypes and ViT features to obtain input features. It mitigates overfitting caused by sample imbalance between novel and base classes.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines