Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task Learning
Xiang Zhang, Wanqing Zhao, Pengyang Li, Ying Liu, Hangzai Luo, Sheng Zhong, Jinye Peng, Jianping Fan
Abstract
Generative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often undermine their effectiveness in downstream visual tasks. This paper introduces the Iterative Self-Training with Class-Aware Text-to-Image Synthesis (IST-CATS) framework, which addresses these challenges by integrating a class-aware text-to-image synthesis (CATS) component with an iterative self-training (IST) strategy. CATS innovatively introduces a class-aware chain approach to generate detailed descriptions. These descriptions act as prompts for a diffusion model, enabling the creation of a diverse of images accompanied by distinguishable objects against the background. The generated images can be easily pseudo-labeled by an unsupervised instance segmentation method, and then noisy pseudo labels can be effectively purified by a novel feature similarity-based filtering mechanism. The generated images underpin our IST, which progressively enhances vision models and refines pseudo labels through self-training and our proposed label filtering strategy (LabFilt). LabFilt meticulously improves the quality of pseudo labels by employing class-adaptive techniques at both the pixel and object levels, ensuring refined pseudo-label accuracy. IST-CATS demonstrates superior performance in object detection and semantic segmentation compared to traditional synthetic and semi/weakly-supervised methods, effectively addressing data collection and annotation challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui et al.NeurIPS 2020 · 755 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
- ST++: Make Self-trainingWork Better for Semi-supervised Semantic SegmentationLihe Yang, Wei Zhuo, Lei Qi, Yinghuan Shi et al.CVPR 2022 · 467 citations
Related papers
- Dataset Diffusion: Diffusion-based Synthetic Data Generation for Pixel-Level Semantic SegmentationQuang Nguyen, Truong Vu, Anh Tran, Khoi NguyenNeurIPS 2023 · 154 citations
- Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic SegmentationDong Zhao, Qi Zang, Shuang Wang, Nicu Sebe et al.ICCV 2025 · 3 citations
- Learning Beyond Vision: Vision-Language Distillation and Edge-Aware Mix Diffusion in Semi-Supervised Semantic SegmentationRui Yang, Yunfei Bai, Yuehua Liu, Xiaomao Li et al.AAAI 2026
- Text-Image Alignment for Diffusion-Based PerceptionNeehar Kondapaneni, Markus Marks, Manuel Knott, Rogério Guimarães et al.CVPR 2024
- DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion ModelsWeijia Wu, Yuzhong Zhao, Hao Chen, Yuchao Gu et al.NeurIPS 2023 · 191 citations
