Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
Xiao Fang, Minhyek Jeon, Shuowen Hu, Zheyang Qin, Shayok Chakraborty, Stanislav Panev, Celso de Melo, Fernando De la Torre
摘要
Detecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data from one geographic region fail to generalize effectively to other areas. Variability in factors such as environmental conditions, urban layouts, road networks, vehicle types, and image acquisition parameters (e.g., resolution, lighting, and angle) leads to domain shifts that degrade model performance. This paper proposes a novel method that uses generative AI to synthesize high-quality aerial images and their labels, improving detector training through data augmentation. Our key contribution is the development of a multi-stage, multi-modal knowledge transfer framework utilizing fine-tuned latent diffusion models (LDMs) to mitigate the distribution gap between the source and target environments. Extensive experiments across diverse aerial imagery domains show consistent performance improvements in AP 50 over supervised learning on source domain data, weakly supervised adaptation methods, unsupervised domain adaptation methods, and open-set object detectors by 4-23%, 6-10%, 7-40%, and more than 50%, respectively. Furthermore, we introduce two newly annotated aerial datasets from New Zealand and Utah to support further research in this field. Project page is available at: https://humansensinglab.github.io/ AGenDA
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned DiffusionZhengkang Xiang, Zizhao Li, Amir Khodabandeh, Kourosh KhoshelhamICCV 2025
- Visual Prototype Conditioned Focal Region Generation for UAV-Based Object DetectionWenhao Li, Zimeng Wu, Yu Wu, Zehua Fu 等CVPR 2026
- Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized DetectionBoyong He, Yuxiang Ji, Qianwen Ye, Zhuoyue Tan 等CVPR 2025
- CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object DetectionGyusam Chang, Wonseok Roh, Sujin Jang, Dongwook Lee 等AAAI 2024 · 被引用 8 次
- Diffusion-Based Source-Biased Model for Single Domain Generalized Object DetectionHan Jiang, Wenfei Yang, Tianzhu Zhang, Yongdong ZhangICCV 2025 · 被引用 2 次
