SCalDA: Semantics-Calibrated and Diffusion-Enhanced Data Augmentation
Shibo Lv, Jianmin Jiang
Abstract
With the rapid development of deep learning, the issue of data scarcity has become increasingly prominent, inspiring emerging interests towards research on data augmentation techniques over recent years. However, our literature survey indicates that existing efforts often suffer from two issues of semantic infidelity, including: (i) visual semantics infidelity, such as visual artifacts, manifold intrusion, and unnatural blending boundaries etc, and (ii) label semantic infidelity, where augmented images do not match the original labels, creating extra label noises. To address these issues, we propose a Semantics Calibrated and Diffusion-Enhanced Augmentation (SCalDA) scheme to achieve accurate semantics calibration across image, label and feature domains. Compared with the existing approaches, our proposed features in precise guidance in label domain, semantics driven synthesis across three domains (image, label and feature), and semantics-aware metric learning. Extensive experiments on multiple datasets demonstrate that SCalDA yields consistent and significant performance improvements for both fine-grained and general classification tasks, validating the effectiveness and broad applicability of the proposed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a638b49-a5d3-495a-a5ac-6fa32a70d6d6Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Beyond One-Hot Labels: Semantic Mixing for Model CalibrationHaoyang Luo, Linwei Tao, Minjing Dong, Chang XuICML 2025
- SAFLEX: Self-Adaptive Augmentation via Feature Label ExtrapolationMucong Ding, Bang An, Yuancheng Xu, Anirudh Satheesh et al.ICLR 2024 · 1 citation
- Tailoring Mixup to Data for CalibrationQuentin Bouniot, Pavlo Mozharovskyi, Florence d'Alché-BucICLR 2025
- KeepAugment: A Simple Information-Preserving Data Augmentation ApproachChengyue Gong, Dilin Wang, Meng Li, Vikas Chandra et al.CVPR 2021
- Quality-Aware Self-Training on Differentiable Synthesis of Rare Relational DataChongsheng Zhang, Yaxin Hou, Ke Chen, Shuang Cao et al.AAAI 2023 · 8 citations
