Salient Concept-Aware Generative Data Augmentation
Tianchen Zhao, Xuanbai Chen, Zhihua Li, Jun Fang, Dongsheng An, Xiang Xu, Zhuowen Tu, Yifan Xing
Abstract
Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73% and 6.5% under conventional and long-tail settings, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e90776b-79a5-4723-a9e4-3fbabb7dadccCited by top-tier papers1
Ask how each one uses itBuilds on60
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- CoRe: Context-Regularized Text Embedding Learning for Text-to-Image PersonalizationFeize Wu, Yun Pang, Junyi Zhang, Lianyu Pang et al.AAAI 2025 · 13 citations
- Decoupled Textual Embeddings for Customized Image GenerationYufei Cai, Yuxiang Wei, Zhilong Ji, Jinfeng Bai et al.AAAI 2024 · 24 citations
- U-VAP: User-specified Visual Appearance Personalization via Decoupled Self AugmentationYou Wu, Kean Liu, Xiaoyue Mi, Fan Tang et al.CVPR 2024 · 6 citations
- AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image GenerationLianyu Pang, Jian Yin, Baoquan Zhao, Feize Wu et al.NeurIPS 2024 · 18 citations
- ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token OptimizationMinseo Kim, Minchan Kwon, Dongyeun Lee, Yunho Jeon et al.CVPR 2026 · 1 citation
