Progressive Transformation Learning for Leveraging Virtual Images in Training
Yi-Ting Shen, Hyungtae Lee, Heesung Kwon, Shuvra S. Bhattacharyya
摘要
To effectively interrogate UAV-based images for detecting objects of interest, such as humans, it is essential to acquire large-scale UAV-based datasets that include human instances with various poses captured from widely varying viewing angles. As a viable alternative to laborious and costly data curation, we introduce Progressive Transformation Learning (PTL), which gradually augments a training dataset by adding transformed virtual images with enhanced realism. Generally, a virtual2real transformation generator in the conditional GAN framework suffers from quality degradation when a large domain gap exists between real and virtual images. To deal with the domain gap, PTL takes a novel approach that progressively iterates the following three steps: 1) select a subset from a pool of virtual images according to the domain gap, 2) transform the selected virtual images to enhance realism, and 3) add the transformed virtual images to the training set while removing them from the pool. In PTL, accurately quantifying the domain gap is critical. To do that, we theoretically demonstrate that the feature representation space of a given object detector can be modeled as a multivariate Gaussian distribution from which the Mahalanobis distance between a virtual object and the Gaussian distribution of each object category in the representation space can be readily computed. Experiments show that PTL results in a substantial performance increase over the baseline, especially in the small data and the cross-domain regime.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingDominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind 等CVPR 2025
- Visual Prototype Conditioned Focal Region Generation for UAV-Based Object DetectionWenhao Li, Zimeng Wu, Yu Wu, Zehua Fu 等CVPR 2026
它引用的顶会 Paper15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- SynFace: Face Recognition with Synthetic DataHaibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li 等ICCV 2021 · 被引用 162 次
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- Label, Verify, Correct: A Simple Few Shot Object Detection MethodPrannay Kaul, Weidi Xie, Andrew ZissermanCVPR 2022 · 被引用 123 次
相关 Paper
- Delving Into Robust Object Detection From Unmanned Aerial Vehicles: A Deep Nuisance Disentanglement ApproachZhenyu Wu, Karthik Suresh, Priya Narayanan, Hongyu Xu 等ICCV 2019 · 被引用 86 次
- Progressive Graph Learning for Open-Set Domain AdaptationYadan Luo, Zijian Wang, Zi Huang, Mahsa BaktashmotlaghICML 2020 · 被引用 114 次
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 被引用 191 次
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionXiao Fang, Minhyek Jeon, Shuowen Hu, Zheyang Qin 等ICCV 2025
- Transformation GAN for Unsupervised Image Synthesis and Representation LearningJiayu Wang, Wengang Zhou, Guo-Jun Qi, Zhongqian Fu 等CVPR 2020
