Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
Jianhao Yuan, Francesco Pinto, Adam Davies, Philip Torr
Abstract
Neural image classifiers are known to undergo severe performance degradation when exposed to inputs that are sampled from environmental conditions that differ from their training data. Given the recent progress in Text-to-Image (T2I) generation, a natural question is how modern T2I generators can be used to simulate arbitrary interventions over such environmental factors in order to augment training data and improve the robustness of downstream classifiers. We experiment across a diverse collection of benchmarks in single domain generalization (SDG) and reducing reliance on spurious features (RRSF), ablating across key dimensions of T2I generation, including interventional prompting strategies, conditioning mechanisms, and post-hoc filtering. Our extensive empirical findings demonstrate that modern T2I generators like Stable Diffusion can indeed be used as a powerful interventional data augmentation mechanism, outperforming previously state-of-the-art data augmentation techniques regardless of how each dimension is configured.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic ObjectChenshuang Zhang, Fei Pan, Junmo Kim, In So Kweon et al.CVPR 2024 · 9 citations
- DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust ClassifiersChandramouli Shama Sastry, Sri Harsha Dumpala, Sageev OoreNeurIPS 2024 · 5 citations
- Distributionally Generative Augmentation for Fair Facial Attribute ClassificationFengda Zhang, Qianpei He, Kun Kuang, Jiashuo Liu et al.CVPR 2024 · 4 citations
- An Analysis of Causal Effect Estimation using Outcome Invariant Data AugmentationUzair Akbar, Niki Kilbertus, Hao Shen, Krikamol Muandet et al.NeurIPS 2025 · 3 citations
- Salient Concept-Aware Generative Data AugmentationTianchen Zhao, Xuanbai Chen, Zhihua Li, Jun Fang et al.NeurIPS 2025 · 2 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Adversarial Domain Prompt Tuning and Generation for Single Domain GeneralizationZhipeng Xu, De Cheng, Xinyang Jiang, Nannan Wang et al.CVPR 2025
- Understanding and Mitigating Copying in Diffusion ModelsGowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping et al.NeurIPS 2023 · 265 citations
- Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet ClonesMert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis KalantidisCVPR 2023
- Your Diffusion Model is Secretly a Zero-Shot ClassifierAlexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown et al.ICCV 2023 · 341 citations
- SafeGuider: Robust and Practical Content Safety Control for Text-to-Image ModelsPeigui Qi, Kunsheng Tang, Wenbo Zhou, Weiming Zhang et al.CCS 2025
