Progressive Transformation Learning for Leveraging Virtual Images in Training
Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya
Abstract
To effectively interrogate UAV-based images for detecting objects of interest, such as humans, it is essential to acquire large-scale UAV-based datasets that include human instances with various poses captured from widely varying viewing angles. As a viable alternative to laborious and costly data curation, we introduce Progressive Transformation Learning (PTL), which gradually augments a training dataset by adding transformed virtual images with enhanced realism. Generally, a virtual2real transformation generator in the conditional GAN framework suffers from quality degradation when a large domain gap exists between real and virtual images. To deal with the domain gap, PTL takes a novel approach that progressively iterates the following three steps: 1) select a subset from a pool of virtual images according to the domain gap, 2) transform the selected virtual images to enhance realism, and 3) add the transformed virtual images to the training set while removing them from the pool. In PTL, accurately quantifying the domain gap is critical. To do that, we theoretically demonstrate that the feature representation space of a given object detector can be modeled as a multivariate Gaussian distribution from which the Mahalanobis distance between a virtual object and the Gaussian distribution of each object category in the representation space can be readily computed. Experiments show that PTL results in a substantial performance increase over the baseline, especially in the small data and the cross-domain regime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02b096f5-0b57-4f1e-8ac4-cf1d41c058f8Cited by top-tier papers2
- LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingDominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind et al.CVPR 2025
- Visual Prototype Conditioned Focal Region Generation for UAV-Based Object DetectionWenhao Li, Zimeng Wu, Yu Wu, Zehua Fu et al.CVPR 2026
Builds on15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- SynFace: Face Recognition with Synthetic DataHaibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li et al.ICCV 2021 · 162 citations
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- Label, Verify, Correct: A Simple Few Shot Object Detection MethodPrannay Kaul, Weidi Xie, Andrew ZissermanCVPR 2022 · 123 citations
Related papers
- Delving Into Robust Object Detection From Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement ApproachZhenyu Wu, Karthik Suresh, Priya Narayanan, Hongyu Xu et al.ICCV 2019 · 86 citations
- Progressive Graph Learning for Open-Set Domain AdaptationYadan Luo, Zijian Wang, Zi Huang, Mahsa BaktashmotlaghICML 2020 · 114 citations
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 191 citations
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionXiao Fang, Minhyek Jeon, Shuowen Hu, Zheyang Qin et al.ICCV 2025
- Transformation GAN for Unsupervised Image Synthesis and Representation LearningJiayu Wang, Wengang Zhou, Guo-Jun Qi, Zhongqian Fu et al.CVPR 2020
