Learning a Dynamic Cross-Modal Network for Multispectral Pedestrian Detection
Jin Xie, Rao Muhammad Anwer, Hisham Cholakkal, Jing Nie, Jiale Cao, Jorma Laaksonen, Fahad Shahbaz Khan
Abstract
Multispectral pedestrian detection that enables continuous (day and night) localization of pedestrians has numerous applications. Existing approaches typically aggregate multispectral features by a simple element-wise operation. However, such a local feature aggregation scheme ignores the rich non-local contextual information. Further, we argue that a local tight correspondence across modalities is desired for multi-modal feature aggregation. To address these issues, we introduce a multispectral pedestrian detection framework that comprises a novel dynamic cross-modal network (DCMNet), which strives to adaptively utilize the local and non-local complementary information between multi-modal features. The proposed DCMNet consists of a local and a non-local feature aggregation module. The local module employs dynamically learned convolutions to capture local relevant information across modalities. On the other hand, the non-local module captures non-local cross-modal information by first projecting features from both modalities into the latent space and then obtaining dynamic latent feature nodes for feature aggregation. Comprehensive experiments are performed on two challenging benchmarks: KAIST and LLVIP. Experiments reveal the benefits of the proposed DCMNet, leading to consistently improved detection performance on diverse detection paradigms and backbones. When using the same backbone, our proposed detector achieves absolute gains of 1.74% and 1.90% over the baseline Cascade RCNN on the KAIST and LLVIP datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a5c5ea91-5ee4-4f2b-bad4-ceff7f742acdCited by top-tier papers1
Ask how each one uses itRelated papers
- Attentive Alignment Network for Multispectral Pedestrian DetectionNuo Chen, Jin Xie, Jing Nie, Jiale Cao et al.ACM MM 2023 · 26 citations
- Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian DetectionLu Zhang, Xiangyu Zhu, Xiangyu Chen, Xu Yang et al.ICCV 2019 · 209 citations
- Causal Mode Multiplexer: A Novel Framework for Unbiased Multispectral Pedestrian DetectionTaeheon Kim, Sebin Shin, Youngjoon Yu, Hak Gu Kim et al.CVPR 2024
- Multispectral Object Detection via Cross-Modal Conflict-Aware LearningXiao He, Chang Tang, Xin Zou, Wei ZhangACM MM 2023 · 84 citations
- Contextually-Guided State Space Fusion for Misaligned Multi-Spectral Object DetectionGuyue Jin, Tianming Zhao, Jiacan Yan, Tian TianACM MM 2025 · 1 citation
