SemanticRT: A Large-Scale Dataset and Method for Robust Semantic Segmentation in Multispectral Images
Wei Ji, Jingjing Li, Cheng Bian, Zhicheng Zhang, Li Cheng
Abstract
Growing interests in multispectral semantic segmentation (MSS) have been witnessed in recent years, thanks to the unique advantages of combining RGB and thermal infrared images to tackle challenging scenarios with adverse conditions. However, unlike traditional RGB-only semantic segmentation, the lack of a large-scale MSS dataset has become a hindrance to the progress of this field. To address this issue, we introduce a SemanticRT dataset - the largest MSS dataset to date, comprising 11,371 high-quality, pixel-level annotated RGB-thermal image pairs. It is 7 times larger than the existing MFNet dataset, and covers a wide variety of challenging scenarios in adverse lighting conditions such as low-light and pitch black. Further, a novel Explicit Complement Modeling (ECM) framework is developed to extract modality-specific information, which is propagated through a robust cross-modal feature encoding and fusion process. Extensive experiments demonstrate the advantages of our approach and dataset over the existing counterparts. Our new dataset may also facilitate further development and evaluation of existing and new MSS algorithms.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers6
- Learning Spectral-Decomposited Tokens for Domain Generalized Semantic SegmentationJingjun Yi, Qi Bi, Hao Zheng, Haolan Zhan et al.ACM MM 2024 · 25 citations
- Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View BenchmarkWei Ji, Jingjing Li, Wenbo Li, Yilin Shen et al.NeurIPS 2024 · 8 citations
- SAM3-I: Segment Anything with InstructionsJingjing Li, Yue Feng, Yuchen Guo, Jincai Huang et al.ACL 2026 · 7 citations
- M-SpecGene: Generalized Foundation Model for RGBT Multispectral VisionKailai Zhou, Fuqiang Yang, Shixian Wang, Bihan Wen et al.ICCV 2025 · 4 citations
- One-shot In-context Part SegmentationZhenqi Dai, Ting Liu, Xingxing Zhang, Yunchao Wei et al.ACM MM 2024 · 1 citation
Related papers
- ABMDRNet: Adaptive-Weighted Bi-Directional Modality Difference Reduction Network for RGB-T Semantic SegmentationQiang Zhang, Shenlu Zhao, Yongjiang Luo, Dingwen Zhang et al.CVPR 2021
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 1 citation
- Robust Multi-Modality Person Re-identificationAihua Zheng, Zi Wang, Zi-Han Chen, Chenglong Li et al.AAAI 2021 · 79 citations
- Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachYun Xiao, Yuhang Wang, Jiandong Jin, Wankang Zhang et al.AAAI 2026
- Edge-Aware Guidance Fusion Network for RGB-Thermal Scene ParsingWujie Zhou, Shaohua Dong, Caie Xu, Yaguan QianAAAI 2022 · 151 citations
