Multimodal Causal Reasoning for UAV Object Detection
Nianxin Li, Mao Ye, Lihua Zhou, Shuaifeng Li, Song Tang, Luping Ji, Ce Zhu
摘要
Unmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited information, complicating detection tasks. To address these challenges, we propose a Multimodel Causal Reasoning framework based on YOLO backbone for UAV Object Detection (MCR-UOD). The key idea is to use the backdoor adjustment to discover the condition-invariant object representation for easy detection. Specifically, the YOLO backbone is first adjusted to incorporate the pre-trained vision-language model. The original category labels are replaced with semantic text prompts, and the detection head is replaced with text-image contrastive learning. Based on this backbone, our method consists of two parts. The first part, named language guided region exploration, discovers the regions with high probability of object existence using text embeddings based on vision-language model such as CLIP. Another part is the backdoor adjustment casual reasoning module, which constructs a confounder dictionary tailored to different imaging conditions to capture global image semantics and derives a prior probability distribution of shooting conditions. During causal inference, we use the confounder dictionary and the prior to intervene on local instance features, disentangling condition variations, and obtaining condition-invariant representations. Experimental results on several public datasets confirm the state-of-the-art performance of our approach. The code, data and models will be released upon publication of this paper.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- Clustered Object Detection in Aerial ImagesFan Yang, Heng Fan, Peng Chu, Erik Blasch 等ICCV 2019 · 被引用 384 次
- Causality Inspired Representation Learning for Domain GeneralizationFangrui Lv, Jian Liang, Shuang Li, Bin Zang 等CVPR 2022 · 被引用 190 次
相关 Paper
- Self-Prompting Analogical Reasoning for UAV Object DetectionNianxin Li, Mao Ye, Lihua Zhou, Song Tang 等AAAI 2025 · 被引用 10 次
- A Causal Marriage between VLM and IRM from Understanding to ReasoningZiliang Chen, Tianang Xiao, jusheng zhang, Yongsen Zheng 等CVPR 2026
- Uni-YOLO: Vision-Language Model-Guided YOLO for Robust and Fast Universal Detection in the Open WorldXudong Wang, Weihong Ren, Xi'ai Chen, Huijie Fan 等ACM MM 2024 · 被引用 6 次
- CLIP2UDA: Making Frozen CLIP Reward Unsupervised Domain Adaptation in 3D Semantic SegmentationYao Wu, Mingwei Xing, Yachao Zhang, Yuan Xie 等ACM MM 2024 · 被引用 12 次
- Representation-Level Counterfactual Calibration for Debiased Zero-Shot RecognitionPei Peng, Ming-Kun Xie, Hang Hao, Tong Jin 等NeurIPS 2025 · 被引用 2 次
