Foreground-Background Separation through Concept Distillation from Generative Image Foundation Models
Mischa Dombrowski, Hadrien Reynaud, Matthew Baugh, Bernhard Kainz
Abstract
Curating datasets for object segmentation is a difficult task. With the advent of large-scale pre-trained generative models, conditional image generation has been given a significant boost in result quality and ease of use. In this paper, we present a novel method that enables the generation of general foreground-background segmentation models from simple textual descriptions, without requiring segmentation labels. We leverage and explore pre-trained latent diffusion models, to automatically generate weak segmentation masks for concepts and objects. The masks are then used to fine-tune the diffusion model on an inpainting task, which enables fine-grained removal of the object, while at the same time providing a synthetic foreground and background dataset. We demonstrate that using this method beats previous methods in both discriminative and generative performance and closes the gap with fully supervised training while requiring no pixel-wise object labels. We show results on the task of segmenting four different objects (humans, dogs, cars, birds) and a use case scenario in medical image analysis. The code is available at https://github.com/MischaD/fobadiffusion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3248b39-c7e6-409a-a3fd-0ea57283f5aeCited by top-tier papers3
- G2D: From Global to Dense Radiography Representation Learning via Vision-Language Pre-trainingChe Liu, Cheng Ouyang, Sibo Cheng, Anand Shah et al.NeurIPS 2024 · 21 citations
- Deciphering 'What' and 'Where' Visual Pathways from Spectral Clustering of Layer-Distributed Neural RepresentationsXiao Zhang, David Yunis, Michael MaireCVPR 2024 · 1 citation
- SCCS: Deep Neural Spectral Clustering for Self-Supervised Subcellular Structure SegmentationJimao Jiang, Diya Sun, Tianbing Wang, Yuru PeiAAAI 2025 · 1 citation
Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Paint by Inpaint: Learning to Add Image Objects by Removing Them FirstNavve Wasserman, Noam Rotstein, Roy Ganz, Ron KimmelCVPR 2025
- Zero-shot spatial layout conditioning for text-to-image diffusion modelsGuillaume Couairon, Marlène Careil, Matthieu Cord, Stéphane Lathuilière et al.ICCV 2023 · 82 citations
- LD-ZNet: A Latent Diffusion Approach for Text-Based Image SegmentationKoutilya PNVR, Bharat Singh, Pallabi Ghosh, Behjat Siddiquie et al.ICCV 2023 · 36 citations
- JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation PromotionHaoyu Wang, Lei Zhang, Wenrui Liu, Dengyang Jiang et al.AAAI 2026
- Tuning-Free Amodal Segmentation via the Occlusion-Free Bias of Inpainting ModelsJae Joong Lee, Bedrich Benes, Raymond A. YehAAAI 2026 · 2 citations
