A Training-free Synthetic Data Selection Method for Semantic Segmentation
Hao Tang, Siyue Yu, Jian Pang, Bingfeng Zhang
Abstract
Training semantic segmenter with synthetic data has been attracting great attention due to its easy accessibility and huge quantities. Most previous methods focused on producing large-scale synthetic image-annotation samples and then training the segmenter with all of them. However, such a solution remains a main challenge in that the poor-quality samples are unavoidable, and using them to train the model will damage the training process. In this paper, we propose a training-free Synthetic Data Selection (SDS) strategy with CLIP to select high-quality samples for building a reliable synthetic dataset. Specifically, given massive synthetic image-annotation pairs, we first design a Perturbation-based CLIP Similarity (PCS) to measure the reliability of synthetic image, thus removing samples with low-quality images. Then we propose a class-balance Annotation Similarity Filter (ASF) by comparing the synthetic annotation with the response of CLIP to remove the samples related to low-quality annotations. The experimental results show that using our method significantly reduces the data size by half, while the trained segmenter achieves higher performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 782b0d32-9241-4620-9f94-d29fc78c7841Cited by top-tier papers3
- Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic SegmentationYunkai Yang, Yudong Zhang, Kunquan Zhang, Jinxiao Zhang et al.CVPR 2026 · 2 citations
- JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation PromotionHaoyu Wang, Lei Zhang, Wenrui Liu, Dengyang Jiang et al.AAAI 2026
- What Makes Synthetic Data Effective in Image SegmentationJinjin Zhang, Xiefan Guo, Yizhou jin, Nan Zhou et al.ICML 2026
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- EditGAN: High-Precision Semantic Image EditingHuan Ling, Karsten Kreis, Daiqing Li, Seung Wook Kim et al.NeurIPS 2021 · 248 citations
Related papers
- ReME: A Data-Centric Framework for Training-Free Open-Vocabulary SegmentationXiwei Xuan, Ziquan Deng, Kwan-Liu MaICCV 2025 · 3 citations
- FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation ModelsLihe Yang, Xiaogang Xu, Bingyi Kang, Yinghuan Shi et al.NeurIPS 2023 · 94 citations
- Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic SegmentationShuo Jin, Siyue Yu, Bingfeng Zhang, Mingjie Sun et al.ICCV 2025 · 3 citations
- CLIPCleaner: Cleaning Noisy Labels with CLIPChen Feng, Georgios Tzimiropoulos, Ioannis PatrasACM MM 2024 · 12 citations
- Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic SegmentationByeongCheol Lee, Hyun Seok Seong, Sangeek Hyun, Gilhan Park et al.CVPR 2026 · 2 citations
