When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
Krzysztof Adamkiewicz, Brian B. Moser, Stanislav Frolov, Tobias Christian Nauen, Federico Raue, Andreas Dengel
摘要
Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as synthetic vision data generators? In this work, we revisit the promise of synthetic data as a scalable substitute for real training sets and uncover a surprising performance regression. We generate large-scale synthetic datasets using state-of-the-art T2I models released between 2022 and 2025, train standard classifiers solely on this synthetic data, and evaluate them on real test data. Despite observable advances in visual fidelity and prompt adherence, classification accuracy on real test data consistently declines with newer T2I models as training data generators. Our analysis reveals a hidden trend: These models collapse to a narrow, aesthetic-centric distribution that undermines diversity and real data distribution coverage. Overall, our findings challenge a growing assumption in vision research, namely that progress in generative realism implies progress in data realism. We thus highlight an urgent need to rethink the capabilities of modern T2I models as reliable training data generators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs BetterScott Geng, Cheng-Yu Hsieh, Vivek Ramanujan, Matthew Wallingford 等NeurIPS 2024 · 被引用 27 次
- The Intricate Dance of Prompt Complexity, Quality, Diversity and Consistency in T2I ModelsXiaofeng Zhang, Aaron C. Courville, Michal Drozdzal, Adriana Romero-SorianoICLR 2026 · 被引用 6 次
- Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet ClonesMert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis KalantidisCVPR 2023
- Scaling Laws of Synthetic Images for Model Training ... for NowLijie Fan, Kaifeng Chen, Dilip Krishnan, Dina Katabi 等CVPR 2024
- Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?Zebin You, Xinyu Zhang, Hanzhong Guo, Jingdong Wang 等CVPR 2025
