FreeGen: Bridging Visual-Linguistic Discrepancies Towards Diffusion-based Pixel-level Data Synthesis
Wenzhuang Wang, Mingcan Ma, Yong Chen, Changqun Xia, Zhenbao Liang, Jia Li
Abstract
Text-to-image diffusion model has inspired research into text-to-data synthesis without human intervention, where spatial attentions correlated with semantic entities in text prompts are primarily interpreted as pseudo-masks. However, these vannila attentions often deliver visual-linguistic discrepancies, in which the associations between image features and entity-level tokens are unstable and divergent, yielding inferior masks for realistic applications, especially in more practical open-vocabulary settings. To tackle this issue, we propose a novel text-guided self-driven generative paradigm, termed FreeGen, which addresses the discrepancies by recalibrating intrinsic visual-linguistic correlations and serves as a real-data-free method to automatically synthesize open-vocabulary pixel-level data for arbitrary entities. Specifically, we first learn an Attention Self-Rectification mechanism to reproject the inherent attention matrices to achieve robust semantic alignment, thereby obtaining class-discriminative masks. A Temporal Fluctuation Factor is present to assess mask quality based on its variation over uniform sampling timesteps, enabling the selection of reliable masks. These masks are then employed as self-supervised signals to support the learning of an Entity-level Grounding Decoder in a self-training manner, thus producing open-vocabulary segmentation results. Extensive experiments show that the existing segmenters trained on FreeGen narrow the performance gap with real data counterparts and remarkably outperform the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 071e2aac-a104-4ad9-a930-ad9b05b7c5feCited by top-tier papers1
Ask how each one uses itBuilds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion ModelsPablo Marcos-Manchón, Roberto Alcover-Couso, Juan C. SanMiguel, Jose M. MartínezCVPR 2024 · 9 citations
- Open-vocabulary Object Segmentation with Diffusion ModelsZiyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang et al.ICCV 2023 · 98 citations
- Segment Anyword: Mask Prompt Inversion for Open-Set Grounded SegmentationZhihua Liu, Amrutha Saseendran, Lei Tong, Xilin He et al.ICML 2025
- SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion ModelsJoon Hyun Park, Kumju Jo, Sungyong BaikAAAI 2025 · 2 citations
- Training-Free Open-Vocabulary Segmentation with Offline Diffusion-Augmented Prototype GenerationLuca Barsellotti, Roberto Amoroso, Marcella Cornia, Lorenzo Baraldi et al.CVPR 2024
