Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
Muzhi Zhu, Yang Liu, Zekai Luo, Chenchen Jing, Hao Chen, Guangkai Xu, Xinlong Wang, Chunhua Shen
Abstract
The Diffusion Model has not only garnered noteworthy achievements in the realm of image generation but has also demonstrated its potential as an effective pretraining method utilizing unlabeled data. Drawing from the extensive potential unveiled by the Diffusion Model in both semantic correspondence and open vocabulary segmentation, our work initiates an investigation into employing the Latent Diffusion Model for Few-shot Semantic Segmentation. Recently, inspired by the in-context learning ability of large language models, Few-shot Semantic Segmentation has evolved into In-context Segmentation tasks, morphing into a crucial element in assessing generalist segmentation models. In this context, we concentrate on Few-shot Semantic Segmentation, establishing a solid foundation for the future development of a Diffusion-based generalist model for segmentation. Our initial focus lies in understanding how to facilitate interaction between the query image and the support image, resulting in the proposal of a KV fusion method within the self-attention framework. Subsequently, we delve deeper into optimizing the infusion of information from the support mask and simultaneously re-evaluating how to provide reasonable supervision from the query mask. Based on our analysis, we establish a simple and effective framework named DiffewS, maximally retaining the original Latent Diffusion Model's generative framework and effectively utilizing the pre-training prior. Experimental results demonstrate that our method significantly outperforms the previous SOTA models in multiple settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f31cc107-c614-4fce-bb93-a6d6388499e2Cited by top-tier papers15
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng et al.NeurIPS 2025 · 45 citations
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System CollaborationHao Zhong, Muzhi Zhu, Zongze Du, Zheng Huang et al.NeurIPS 2025 · 40 citations
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsZhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou et al.CVPR 2026 · 36 citations
- A Simple Image Segmentation Framework via In-Context ExamplesYang Liu, Chenchen Jing, Hengtao Li, Muzhi Zhu et al.NeurIPS 2024 · 29 citations
- INSID3: Training-Free In-Context Segmentation with DINOv3Claudia Cuttano, Gabriele Trivigno, Christoph Reich, Daniel Cremers et al.CVPR 2026 · 13 citations
Builds on41
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
Related papers
- Few-shot Semantic Segmentation via Perceptual Attention and Spatial ControlGuangchen Shi, Wei Zhu, Yirui Wu, Danhuai Zhao et al.ACM MM 2024 · 3 citations
- DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationAmin Karimi, Charalambos PoullisCVPR 2025
- Explore In-Context Segmentation via Latent Diffusion ModelsChaoyang Wang, Xiangtai Li, Henghui Ding, Lu Qi et al.AAAI 2025 · 17 citations
- Label-Efficient Semantic Segmentation with Diffusion ModelsDmitry Baranchuk, Andrey Voynov, Ivan Rubachev, Valentin Khrulkov et al.ICLR 2022 · 700 citations
- SLiMe: Segment Like MeAliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi, Ali Mahdavi-Amiri et al.ICLR 2024 · 47 citations
