Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
Junjiao Tian, Lavisha Aggarwal, Andrea Colaco, Zsolt Kira, Mar González-Franco
摘要
Producing quality segmentation masks for images is a fundamental problem in computer vision. Recent research has explored large-scale supervised training to enable zero- shot transfer segmentation on virtually any image style and unsupervised training to enable segmentation without dense annotations. However, constructing a model capable of segmenting anything in a zero-shot manner without any anno-tations is still challenging. In this paper, we propose to uti-lize the self-attention layers in stable diffusion models to achieve this goal because the pre-trained stable diffusion model has learned inherent concepts of objects within its attention layers. Specifically, we introduce a simple yet ef-fective iterative merging process based on measuring KL divergence among attention maps to merge them into valid segmentation masks. The proposed method does not re-quire any training or language dependency to extract qual-ity segmentation for any images. On COCO-Stuff-27, our method surpasses the prior unsupervised zero-shot trans-fer SOTA method by an absolute 26% in pixel accuracy and 17% in mean IoU. The project page is at https://sites.google.com/view/diffseg/home.<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Georgia Institute of Technology
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper60
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng 等NeurIPS 2024 · 被引用 291 次
- Augmented Object Intelligence with XR-ObjectsMustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du 等UIST 2024 · 被引用 64 次
- SLiMe: Segment Like MeAliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi, Ali Mahdavi-Amiri 等ICLR 2024 · 被引用 47 次
- Interpretable Diffusion via Information DecompositionXianghao Kong, Ollie Liu, Han Li, Dani Yogatama 等ICLR 2024 · 被引用 37 次
- DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized CutPaul Couairon, Mustafa Shukor, Jean-Emmanuel Haugeard, Matthieu Cord 等NeurIPS 2024 · 被引用 32 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
相关 Paper
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion ModelsJoon Hyun Park, Kumju Jo, Sungyong BaikAAAI 2025 · 被引用 2 次
- DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion ModelsWeijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou 等ICCV 2023 · 被引用 198 次
- Few-shot Semantic Segmentation via Perceptual Attention and Spatial ControlGuangchen Shi, Wei Zhu, Yirui Wu, Danhuai Zhao 等ACM MM 2024 · 被引用 3 次
- EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion ModelsKoichi Namekata, Amirmojtaba Sabour, Sanja Fidler, Seung Wook KimICLR 2024 · 被引用 41 次
- Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive SegmentationMarkus Karmann, Onay UrfaliogluCVPR 2025
