Rethinking Multi-Modal Object Detection From the Perspective of Mono-Modality Feature Learning
Tianyi Zhao, Boyang Liu, Yanglei Gao, Yiming Sun, Maoxun Yuan, Xingxing Wei
摘要
Multi-Modal Object Detection (MMOD), due to its stronger adaptability to various complex environments, has been widely applied in various applications. Extensive research is dedicated to the RGB-IR object detection, primarily focusing on how to integrate complementary features from RGB-IR modalities. However, they neglect the monomodality insufficient learning problem, which arises from decreased feature extraction capability in multi-modal joint learning. This leads to a prevalent but unreasonable phe-nomenon-Fusion Degradation, which hinders the performance improvement of the MMOD model. Motivated by this, in this paper, we introduce linear probing evaluation to the multi-modal detectors and rethink the multimodal object detection task from the mono-modality learning perspective. Therefore, we construct a novel framework called -LIF, which consists of the Mono-Modality Distillation () method and the Local Illuminationaware Fusion (LIF) module. The D-LIF framework facilitates the sufficient learning of mono-modality during multi-modal joint training and explores a lightweight yet effective feature fusion manner to achieve superior object detection performance. Extensive experiments conducted on three MMOD datasets demonstrate that our D-LIF effectively mitigates the Fusion Degradation phenomenon and outperforms the previous SOTA detectors. The codes are available at https://github.com/Zhao-Tian-yi/M2D-LIF.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Tri-Modal Fusion Transformers for UAV-based Object DetectionCraig Iaboni, Pramod AbichandaniCVPR 2026
- Infrared-Privileged UAV Detection via Cross-Modal Vector-QuantizationZhibo Lou, Ruijie Zhang, Zeyu Luo, Qianxi Cao 等AAAI 2026
- DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object DetectionChaolang Li, Pengwen Dai, Jingyu Li, Siyuan Yao 等CVPR 2026
- UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV DetectionShenghui Huang, Menghao Hu, Longkun Zou, Hongyu Chi 等CVPR 2026
它引用的顶会 Paper15
- SimAM: A Simple, Parameter-Free Attention Module for Convolutional Neural NetworksLingxiao Yang, Ru-Yuan Zhang, Lida Li, Xiaohua XieICML 2021 · 被引用 1,593 次
- Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object DetectionJinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu 等CVPR 2022 · 被引用 929 次
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan 等ICCV 2021 · 被引用 432 次
- Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian DetectionLu Zhang, Xiangyu Zhu, Xiangyu Chen, Xu Yang 等ICCV 2019 · 被引用 209 次
- PKD: General Distillation Framework for Object Detectors via Pearson Correlation CoefficientWeihan Cao, Yifan Zhang, Jianfei Gao, Anda Cheng 等NeurIPS 2022 · 被引用 147 次
相关 Paper
- Distribution-Aligned Multimodal Fusion for Robust Object DetectionXiaohui Hao, Yanglin Pu, Yongjun Wang, Rui SheCVPR 2026
- Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionYanfeng Liu, Lefei ZhangACM MM 2025 · 被引用 2 次
- MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionGuibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang 等ACM MM 2020 · 被引用 53 次
- Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image FusionBing Cao, Yiming Sun, Pengfei Zhu, Qinghua HuICCV 2023 · 被引用 110 次
- Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object TrackingShilei Wang, Pujian Lai, Dong Gao, Jifeng Ning 等AAAI 2026
