Open-vocabulary Object Segmentation with Diffusion Models
Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, Weidi Xie
Abstract
The goal of this paper is to extract the visual-language correspondence from a pre-trained text-to-image diffusion model, in the form of segmentation map, i.e., simultaneously generating images and segmentation masks for the corresponding visual entities described in the text prompt. We make the following contributions: (i) we pair the existing Stable Diffusion model with a novel grounding module, that can be trained to align the visual and textual embedding space of the diffusion model with only a small number of object categories; (ii) we establish an automatic pipeline for constructing a dataset, that consists of image, segmentation mask, text prompt triplets, to train the proposed grounding module; (iii) we evaluate the performance of open-vocabulary grounding on images generated from the text-to-image diffusion model and show that the module can well segment the objects of categories beyond seen ones at training time, as shown in Fig. 1; (iv) we adopt the augmented diffusion model to build a synthetic semantic segmentation dataset, and show that, training a standard segmentation model on such dataset demonstrates competitive performance on the zero-shot segmentation (ZS3) benchmark, which opens up new opportunities for adopting the powerful diffusion model for discriminative tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cddc55ab-9db4-469a-94b1-d4e27473abf4Cited by top-tier papers48
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng et al.NeurIPS 2024 · 291 citations
- GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data GenerationKai Chen, Enze Xie, Zhe Chen, Yibo Wang et al.ICLR 2024 · 60 citations
- Boosting Generative Image Modeling via Joint Image-Feature SynthesisTheodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris et al.NeurIPS 2025 · 47 citations
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
- UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow ModelsGuanlong Jiao, Biqing Huang, Kuan-Chieh Wang, Renjie LiaoICLR 2026 · 42 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- FreeGen: Bridging Visual-Linguistic Discrepancies Towards Diffusion-based Pixel-level Data SynthesisWenzhuang Wang, Mingcan Ma, Yong Chen, Changqun Xia et al.AAAI 2025 · 1 citation
- Zero-shot spatial layout conditioning for text-to-image diffusion modelsGuillaume Couairon, Marlène Careil, Matthieu Cord, Stéphane Lathuilière et al.ICCV 2023 · 82 citations
- DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion ModelsWeijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou et al.ICCV 2023 · 198 citations
- Dataset Diffusion: Diffusion-based Synthetic Data Generation for Pixel-Level Semantic SegmentationQuang Nguyen, Truong Vu, Anh Tran, Khoi NguyenNeurIPS 2023 · 154 citations
- Text-to-Image Diffusion Models are Zero Shot ClassifiersKevin Clark, Priyank JainiNeurIPS 2023 · 192 citations
