Tri-Modal Fusion Transformers for UAV-based Object Detection
Craig Iaboni, Pramod Abichandani
摘要
Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave infrared (LWIR) sensing preserves contrast in low light, and event cameras retain microsecond-level temporal edges, but integrating all three modalities in a unified detector has not been systematically studied. We present a tri-modal framework that processes RGB, thermal, and event data with a dual-stream hierarchical vision transformer. At selected encoder depths, a Modality-Aware Gated Exchange (MAGE) applies intersensor channel and spatial gating, and a Bidirectional Token Exchange (BiTE) module performs bidirectional tokenlevel attention with depthwise-pointwise refinement, producing resolution-preserving fused maps for a standard feature pyramid and two-stage detector.
We introduce a 10,489-frame UAV dataset with synchronized and pre-aligned RGB-thermal-event streams and 24,223 annotated vehicles across day and night flights. Through 61 controlled ablations, we evaluate fusion placement, mechanism (baseline MAGE+BiTE, CSSA, GAFF), modality subsets, and backbone capacity. Tri-modal fusion improves over all dual-modal baselines, with fusion depth having a significant effect and a lightweight CSSA variant recovering most of the benefit at minimal cost. This work provides the first systematic benchmark and modular backbone for tri-modal UAV-based object detection. Code and dataset are available at https://github.com/ radlab-sketch/trimodal-uav-det.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu 等CVPR 2022 · 被引用 929 次
相关 Paper
- Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachYun Xiao, Yuhang Wang, Jiandong Jin, Wankang Zhang 等AAAI 2026
- TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationZeyu Wang, Fabien Colonnier, Jinghong Zheng, Jyotibdha Acharya 等ACM MM 2023 · 被引用 28 次
- IGIANet: Illumination Guided Implicit Alignment Network for Infrared-Visible UAV DetectionXiangqi Chen, Dawei Zhang, Li Zhao, Chengzhuan Yang 等AAAI 2026
- CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkWentao Wu, Xiao Wang, Chenglong Li, Bo Jiang 等ACM MM 2025 · 被引用 2 次
- Robust Pedestrian Detection with Uncertain ModalityQian Bie, Xiao Wang, Bin Yang, Zhixi Yu 等AAAI 2026
