LD-ZNet: A Latent Diffusion Approach for Text-Based Image Segmentation
Koutilya PNVR, Bharat Singh, Pallabi Ghosh, Behjat Siddiquie, David Jacobs
摘要
Large-scale pre-training tasks like image classification, captioning, or self-supervised techniques do not incentivize learning the semantic boundaries of objects. However, recent generative foundation models built using text-based latent diffusion techniques may learn semantic boundaries. This is because they have to synthesize intricate details about all objects in an image based on a text description. Therefore, we present a technique for segmenting real and AI-generated images using latent diffusion models (LDMs) trained on internet-scale datasets. First, we show that the latent space of LDMs (z-space) is a better input representation compared to other feature representations like RGB images or CLIP encodings for text-based image segmentation. By training the segmentation models on the latent z-space, which creates a compressed representation across several domains like different forms of art, cartoons, illustrations, and photographs, we are also able to bridge the domain gap between real and AI-generated images. We show that the internal features of LDMs contain rich semantic information and present a technique in the form of LD-ZNet to further boost the performance of text-based segmentation. Overall, we show up to 6% improvement over standard baselines for text-to-image segmentation on natural images. For AI-generated imagery, we show close to 20% improvement compared to state-of-the-art techniques. The project is available at https://koutilya-pnvr.github.io/LD-ZNet/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Dataset Diffusion: Diffusion-based Synthetic Data Generation for Pixel-Level Semantic SegmentationQuang Nguyen, Truong Vu, Anh Tran, Khoi NguyenNeurIPS 2023 · 被引用 154 次
- Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion ModelsPablo Marcos-Manchón, Roberto Alcover-Couso, Juan C. SanMiguel, Jose M. MartínezCVPR 2024 · 被引用 9 次
- RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing ImagesKe Li, Di Wang, Ting Wang, Fuyu Dong 等AAAI 2026 · 被引用 7 次
- Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging ScenesHaonan Wang, Hanyu Zhou, Haoyue Liu, Luxin YanNeurIPS 2025 · 被引用 4 次
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion ModelsJoon Hyun Park, Kumju Jo, Sungyong BaikAAAI 2025 · 被引用 2 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Foreground-Background Separation through Concept Distillation from Generative Image Foundation ModelsMischa Dombrowski, Hadrien Reynaud, Matthew Baugh, Bernhard KainzICCV 2023 · 被引用 9 次
- Explore In-Context Segmentation via Latent Diffusion ModelsChaoyang Wang, Xiangtai Li, Henghui Ding, Lu Qi 等AAAI 2025 · 被引用 17 次
- LDMol: A Text-to-Molecule Diffusion Model with Structurally Informative Latent Space Surpasses AR ModelsJinho Chang, Jong Chul YeICML 2025
- Segment Any-Quality Images with Generative Latent Space EnhancementGuangqian Guo, Yong Guo, Xuehui Yu, Wenbo Li 等CVPR 2025
- Zero-shot spatial layout conditioning for text-to-image diffusion modelsGuillaume Couairon, Marlène Careil, Matthieu Cord, Stéphane Lathuilière 等ICCV 2023 · 被引用 82 次
