Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition Cues
Chen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi, Xingyue Liu, Tianpeng Liu, Zhen Liu, Yongxiang Liu, Ping Zhong
摘要
Unmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of high-quality dataset. However, the existing dataset struggles to fully capture real-world complexity for limited imaging conditions. To this end, we introduce a high-diversity dataset ATR-UMOD covering varying scenarios, spanning altitudes from 80m to 300m, angles from 0° to 75°, and all-day, all-year time variations in rich weather and illumination conditions. Moreover, each RGB-IR image pair is annotated with 6 condition attributes, offering valuable high-level contextual information. To meet the challenge raised by such diverse conditions, we propose a novel prompt-guided condition-aware dynamic fusion (PCDF) to adaptively reassign multimodal contributions by leveraging annotated condition cues. By encoding imaging conditions as text prompts, PCDF effectively models the relationship between conditions and multimodal contributions through a task-specific soft-gating transformation. A prompt-guided condition-decoupling module further ensures the availability in practice without condition annotations. Experiments on ATR-UMOD dataset reveal the effectiveness of PCDF.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Oriented R-CNN for Object DetectionXingxing Xie, Gong Cheng, Jiabao Wang, Xiwen Yao 等ICCV 2021 · 被引用 1,070 次
- Presence-Only Geographical Priors for Fine-Grained Image ClassificationOisin Mac Aodha, Elijah Cole, Pietro PeronaICCV 2019 · 被引用 206 次
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu 等ICML 2023 · 被引用 143 次
- Multispectral Object Detection via Cross-Modal Conflict-Aware LearningXiao He, Chang Tang, Xin Zou, Wei ZhangACM MM 2023 · 被引用 84 次
相关 Paper
- Infrared-Privileged UAV Detection via Cross-Modal Vector-QuantizationZhibo Lou, Ruijie Zhang, Zeyu Luo, Qianxi Cao 等AAAI 2026
- IGIANet: Illumination Guided Implicit Alignment Network for Infrared-Visible UAV DetectionXiangqi Chen, Dawei Zhang, Li Zhao, Chengzhuan Yang 等AAAI 2026
- CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image FusionYiming Sun, Yuan Ruan, Qinghua Hu, Pengfei ZhuAAAI 2026
- DDFD: Diffusion-Based Denoising Fusion for Object Detection in Infrared-Visible ImagesMin Dang, Gang Liu, Jingqi Zhao, Adams Wai-Kin Kong 等ACM MM 2025 · 被引用 3 次
- Dispel Darkness for Better Fusion: A Controllable Visual Enhancer Based on Cross-Modal Conditional Adversarial LearningHao Zhang, Linfeng Tang, Xinyu Xiang, Xuhui Zuo 等CVPR 2024 · 被引用 21 次
