Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. Little
Abstract
The emergence of attention-based transformer models has led to their extensive use in various tasks, due to their superior generalization and transfer properties. Recent research has demonstrated that such models, when prompted appropriately, are excellent for few-shot inference. However, such techniques are under-explored for dense prediction tasks like semantic segmentation. In this work, we examine the effectiveness of prompting a transformerdecoder with learned visual prompts for the generalized few-shot segmentation (GFSS) task. Our goal is to achieve strong performance not only on novel categories with limited examples, but also to retain performance on base categories. We propose an approach to learn visual prompts with limited examples. These learned visual prompts are used to prompt a multiscale transformer decoder to facilitate accurate dense predictions. Additionally, we introduce a unidirectional causal attention mechanism between the novel prompts, learned with limited examples, and the base prompts, learned with abundant data. This mechanism enriches the novel prompts without deteriorating the base class performance. Overall, this form of prompting helps us achieve state-of-the-art performance for GFSS on two different benchmark datasets: COCO-20 i and Pascal-5 i , without the need for test-time optimization (or transduction). Furthermore, test-time optimization leveraging unlabelled test data can be used to improve the prompts, which we refer to as transductive prompt tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot SegmentationShi-Feng Peng, Guolei Sun, Yong Li, Hongsong Wang et al.AAAI 2025 · 7 citations
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
- DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic SegmentationZiyu Zhao, Xiaoguang Li, Lingjia Shi, Nasrin Imanpour et al.CVPR 2025
- Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language ModelZhaochong An, Guolei Sun, Yun Liu, Runjia Li et al.CVPR 2025
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?Tilemachos Aravanis, Vladan Stojnic, Bill Psomas, Nikos Komodakis et al.CVPR 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- MM-Prompt: Multi-modality and Multi-granularity Prompts for Few-Shot SegmentationHang Xiong, Runmin Cong, Jinpeng Chen, Chen Zhang et al.ACM MM 2025
- DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationAmin Karimi, Charalambos PoullisCVPR 2025
- FS-DETR: Few-Shot DEtection TRansformer with prompting and without re-trainingAdrian Bulat, Ricardo Guerrero, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 61 citations
- Feature-Proxy Transformer for Few-Shot SegmentationJian-Wei Zhang, Yifan Sun, Yi Yang, Wei ChenNeurIPS 2022 · 105 citations
- A Surprisingly Simple Approach to Generalized Few-Shot Semantic SegmentationTomoya Sakai, Haoxiang Qiu, Takayuki Katsuki, Daiki Kimura et al.NeurIPS 2024 · 7 citations
