MMVIP: A Visible-infrared Paired Dataset for Multi-weather Marine Vision
Yunpeng Yin, Lihan Wang, Zhaoshen He, Xinqiang He, Xingming Liao, Zhuowei Wang, Lianglun Cheng
Abstract
Maritime multimodal vision faces significant challenges due to the complexity and variability of oceanic weather and environmental conditions. While modern vessels are commonly equipped with visible and infrared imaging systems, the complementary nature of these modalities fundamentally depends on accurate cross-modal registration. However, the absence of paired visible-infrared datasets that realistically capture diverse maritime scenarios has severely hindered progress in this field. To overcome this limitation, we present MMVIP, the first large-scale visible-infrared maritime vision dataset covering a wide spectrum of weather conditions and sea states. The dataset contains 128,100 images and 50 video sequences with precise spatial-temporal alignment. Comprehensive evaluations across image registration, fusion, maritime object detection, and cross-modal image translation tasks demonstrate the dataset's effectiveness and challenge. Furthermore, MMVIP establishes a new benchmark for advancing multimodal maritime perception. The dataset and corresponding benchmarks are publicly available at: https: //github.com/yyppptjr/MMVIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- Multi-interactive Feature Learning and a Full-time Multi-modality Benchmark for Image Fusion and SegmentationJinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma et al.ICCV 2023 · 287 citations
- Equivariant Multi-Modality Image FusionZixiang Zhao, Haowen Bai, Jiangshe Zhang, Yulun Zhang et al.CVPR 2024 · 155 citations
- Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like SpeedYifan Wang, Xingyi He, Sida Peng, Dongli Tan et al.CVPR 2024 · 126 citations
- Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionXunpeng Yi, Han Xu, Hao Zhang, Linfeng Tang et al.CVPR 2024 · 121 citations
Related papers
- Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachYun Xiao, Yuhang Wang, Jiandong Jin, Wankang Zhang et al.AAAI 2026
- Cross-Modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and MethodHan Wang, Shengyang Li, Jian Yang, Yuxuan Liu et al.ICCV 2025 · 10 citations
- LaRS: A Diverse Panoptic Maritime Obstacle Detection Dataset and BenchmarkLojze Zust, Janez Pers, Matej KristanICCV 2023 · 44 citations
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 1 citation
- Cross-Modal Object Tracking: Modality-Aware Representations and a Unified BenchmarkChenglong Li, Tianhao Zhu, Lei Liu, Xiaonan Si et al.AAAI 2022 · 12 citations
