Infrared-Privileged UAV Detection via Cross-Modal Vector-Quantization
Zhibo Lou, Ruijie Zhang, Zeyu Luo, Qianxi Cao, Feng Qian, Junjie Chen, Yuming Fang
摘要
RGB and infrared images has shown remarkable robustness for object detection based on unmanned aerial vehicles (UAV). However, the primitive RGB and infrared (IR) images are inevitably misaligned due to the device gap between RGB and infrared cameras. Most existing methods rely on manually filtered and aligned images, and thus are limited in real-world application. Some recent methods tend to directly learn from misaligned images, which only weakly benefit from the multi-modality and may be misled by dramatically misaligned IR images. Considering that the manually aligned images are available during training while unavailable in inference, we explore a new learning paradigm using the IR modality as privileged information. In the training stage, our model learns to hallucinate the complementary knowledge in IR modality based on RGB modality. In inference, our model could hallucinate the complementary IR modality to facilitate UAV detection. Specifically, we propose to quantize the IR features and hallucinate the codebook-indices based on RGB features, which is more effective and robust than directly hallucinating features. In addition, we propose to hierarchically hallucinate multi-scale codebook-indices, which could further improve the hallucinating quality. Experiments on DroneVehicle and VisDrone datasets demonstrate the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 被引用 730 次
- Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian DetectionLu Zhang, Xiangyu Zhu, Xiangyu Chen, Xu Yang 等ICCV 2019 · 被引用 209 次
- Multispectral Object Detection via Cross-Modal Conflict-Aware LearningXiao He, Chang Tang, Xin Zou, Wei ZhangACM MM 2023 · 被引用 84 次
- Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged ObjectsJian Hu, Jiayi Lin, Shaogang Gong, Weitong CaiAAAI 2024 · 被引用 64 次
相关 Paper
- IGIANet: Illumination Guided Implicit Alignment Network for Infrared-Visible UAV DetectionXiangqi Chen, Dawei Zhang, Li Zhao, Chengzhuan Yang 等AAAI 2026
- Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object DetectionChen Chen, Jiahao Qi, Xingyue Liu, Kangcheng Bin 等CVPR 2024
- Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachYun Xiao, Yuhang Wang, Jiandong Jin, Wankang Zhang 等AAAI 2026
- Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition CuesChen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi 等ICCV 2025 · 被引用 8 次
- CDUPatch: Color-Driven Universal Adversarial Patch Attack for Dual-Modal Visible-Infrared DetectorsJiahuan Long, Wen Yao, Tingsong Jiang, Jiacheng Hou 等ACM MM 2025 · 被引用 7 次
