OntoAug: Rethinking Generative Data Augmentation via Ontology Guidance
Shuo Wang, Zhichuan Wang, Jun Luo
Abstract
Generative data augmentation has created new opportunities for improving image recognition models. Despite these advances, most existing augmentation methods process images holistically, without accounting for the uneven distribution of discriminative information in classification tasks, where foreground subjects typically contain richer category-relevant signals than background context. Ignoring this imbalance can introduce unintended semantic shifts in generated samples, thereby weakening the model's ability to capture the intrinsic ontology of the target object. In fact, human perception operates in a similar manner: subject identity remains stable, while contextual variations are naturally tolerated as long as overall coherence is preserved. Inspired by this observation, we propose OntoAug, an ontology-oriented data augmentation framework that explicitly distinguishes between subject and environment. On-toAug separates foreground and background through structured layout control and guides diffusion models to generate samples with consistent subjects and diverse contextual variations. By explicitly introducing foreground masks and diverse background prompts during the generation process, the framework enhances both semantic fidelity and diversity in the synthesized data. Extensive experiments show that OntoAug achieves strong performance across multiple tasks, including image classification, few-shot learning, weakly supervised object localization (WSOL), and large vision-language model reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44c14171-430d-47f8-b9fa-1f463d2895dcBuilds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- DRMix: Decomposition-Recomposition Data Augmentation with Diffusion ModelShuo Wang, Zhichuan Wang, Yanmin Chen, Mengyao Zhou et al.ACM MM 2025
- Multi-Perspective Data Augmentation for Few-shot Object DetectionAnh-Khoa Nguyen Vu, Quoc-Truong Truong, Vinh-Tiep Nguyen, Thanh Duc Ngo et al.ICLR 2025
- Effective Data Augmentation With Diffusion ModelsBrandon Trabucco, Kyle Doherty, Max Gurinas, Ruslan SalakhutdinovICLR 2024 · 380 citations
- Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image ClassificationBohan Li, Xiao Xu, Xinghao Wang, Yutai Hou et al.AAAI 2024 · 28 citations
- Advancing Fine-Grained Classification by Structure and Subject Preserving AugmentationEyal Michaeli, Ohad FriedNeurIPS 2024 · 19 citations
