EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models
Koichi Namekata, Amirmojtaba Sabour, Sanja Fidler, Seung Wook Kim
摘要
Diffusion models have recently received increasing research attention for their remarkable transfer abilities in semantic segmentation tasks. However, generating fine-grained segmentation masks with diffusion models often requires additional training on annotated datasets, leaving it unclear to what extent pre-trained diffusion models alone understand the semantic relations of their generated images. To address this question, we leverage the semantic knowledge extracted from Stable Diffusion (SD) and aim to develop an image segmentor capable of generating fine-grained segmentation maps without any additional training. The primary difficulty stems from the fact that semantically meaningful feature maps typically exist only in the spatially lower-dimensional layers, which poses a challenge in directly extracting pixel-level semantic relations from these feature maps. To overcome this issue, our framework identifies semantic correspondences between image pixels and spatial locations of low-dimensional feature maps by exploiting SD's generation process and utilizes them for constructing image-resolution segmentation maps. In extensive experiments, the produced segmentation maps are demonstrated to be well delineated and capture detailed parts of the images, indicating the existence of highly accurate pixel-level semantic knowledge in diffusion models. Project page: https://kmcode1.github.io/Projects/ EmerDiff/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingYunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert 等NeurIPS 2024 · 被引用 56 次
- DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized CutPaul Couairon, Mustafa Shukor, Jean-Emmanuel Haugeard, Matthieu Cord 等NeurIPS 2024 · 被引用 32 次
- Exploring the Underwater World Segmentation without Extra TrainingBingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao 等CVPR 2026 · 被引用 18 次
- Zero-shot Depth Completion via Test-time Alignment with Affine-invariant Depth PriorLee Hyoseok, Kyeong Seon Kim, Byung-Ki Kwon, Tae-Hyun OhAAAI 2025 · 被引用 11 次
- Instruct Where the Model Fails: Generative Data Augmentation via Guided Self-contrastive Fine-tuningWeijian Ma, Ruoxin Chen, Ke-Yue Zhang, Shuang Wu 等AAAI 2025 · 被引用 8 次
它引用的顶会 Paper49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion ModelsJoon Hyun Park, Kumju Jo, Sungyong BaikAAAI 2025 · 被引用 2 次
- Label-Efficient Semantic Segmentation with Diffusion ModelsDmitry Baranchuk, Andrey Voynov, Ivan Rubachev, Valentin Khrulkov 等ICLR 2022 · 被引用 700 次
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersChaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon 等NeurIPS 2025 · 被引用 6 次
- Dataset Diffusion: Diffusion-based Synthetic Data Generation for Pixel-Level Semantic SegmentationQuang Nguyen, Truong Vu, Anh Tran, Khoi NguyenNeurIPS 2023 · 被引用 154 次
- Few-shot Semantic Segmentation via Perceptual Attention and Spatial ControlGuangchen Shi, Wei Zhu, Yirui Wu, Danhuai Zhao 等ACM MM 2024 · 被引用 3 次
