Multi-Perspective Data Augmentation for Few-shot Object Detection
Anh-Khoa Nguyen Vu, Quoc-Truong Truong, Vinh-Tiep Nguyen, Thanh Duc Ngo, Thanh-Toan Do, Tam V. Nguyen
Abstract
Recent few-shot object detection (FSOD) methods have focused on augmenting synthetic samples for novel classes, show promising results to the rise of diffusion models. However, the diversity of such datasets is often limited in representativeness because they lack awareness of typical and hard samples, especially in the context of foreground and background relationships. To tackle this issue, we propose a Multi-Perspective Data Augmentation (MPAD) framework. In terms of foreground-foreground relationships, we propose in-context learning for object synthesis (ICOS) with bounding box adjustments to enhance the detail and spatial information of synthetic samples. Inspired by the large margin principle, support samples play a vital role in defining class boundaries. Therefore, we design a Harmonic Prompt Aggregation Scheduler (HPAS) to mix prompt embeddings at each time step of the generation process in diffusion models, producing hard novel samples. For foreground-background relationships, we introduce a Background Proposal method (BAP) to sample typical and hard backgrounds. Extensive experiments on multiple FSOD benchmarks demonstrate the effectiveness of our approach. Our framework significantly outperforms traditional methods, achieving an average increase of 17.5% in nAP50 over the baseline on PASCAL VOC. Code is available at github.com/nvakhoa/MPAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency LearningBozhao Li, Shaocong Wu, Tong Shao, Senqiao Yang et al.CVPR 2026
- Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCILRuitao Wu, Yifan Zhao, Guangyao Chen, Jia LiNeurIPS 2025
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- OntoAug: Rethinking Generative Data Augmentation via Ontology GuidanceShuo Wang, Zhichuan Wang, Jun LuoCVPR 2026
- Fine-Grained Prototypes Distillation for Few-Shot Object DetectionZichen Wang, Bo Yang, Haonan Yue, Zhenghao MaAAAI 2024 · 55 citations
- MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic SegmentationYuqi Lin, Hao Zhang, Wenqi Shao, Shiqu Liu et al.CVPR 2026
- ODGEN: Domain-specific Object Detection Data Generation with Diffusion ModelsJingyuan Zhu, Shiyu Li, Yuxuan Liu, Jian Yuan et al.NeurIPS 2024 · 32 citations
- SNIDA: Unlocking Few-Shot Object Detection with Non-Linear Semantic Decoupling AugmentationYanjie Wang, Xu Zou, Luxin Yan, Sheng Zhong et al.CVPR 2024 · 22 citations
