Delivering Arbitrary-Modal Semantic Segmentation
Jiaming Zhang, Ruiping Liu, Hao Shi, Kailun Yang, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, Rainer Stiefelhagen
摘要
Multimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the DELIVER arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we provide this dataset in four severe weather conditions as well as five sensor failure cases to exploit modal complementarity and resolve partial outages. To make this possible, we present the arbitrary cross-modal segmentation model CMNEXT. It encompasses a Self-Query Hub (SQ-Hub) designed to extract effective information from any modality for subsequent fusion with the RGB representation and adds only negligible amounts of parameters (∼0.01M ) per additional modality. On top, to efficiently and flexibly harvest discriminative cues from the auxiliary modalities, we introduce the simple Parallel Pooling Mixer (PPX). With extensive experiments on a total of six benchmarks, our CMNEXT achieves state-of-the-art performance on the DELIVER, KITTI-360, MFNet, NYU Depth V2, UrbanLF, and MCubeS datasets, allowing to scale from 1 to 81 modalities. On the freshly collected DELIVER, the quad-modal CMNEXT reaches up to 66.30% in mIoU with a +9.10% gain as compared to the mono-modal baseline. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationBowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu 等ICLR 2024 · 被引用 110 次
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu 等CVPR 2024 · 被引用 78 次
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu 等ICML 2024 · 被引用 64 次
- CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic SegmentationRuihao Xia, Chaoqiang Zhao, Meng Zheng, Ziyan Wu 等ICCV 2023 · 被引用 54 次
- Rethinking Reverse Distillation for Multi-Modal Anomaly DetectionZhihao Gu, Jiangning Zhang, Liang Liu, Xu Chen 等AAAI 2024 · 被引用 51 次
它引用的顶会 Paper43
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
相关 Paper
- MixPrompt: Efficient Mixed Prompting for Multimodal Semantic SegmentationZhiwei Hao, Zhongyu Xiao, Jianyuan Guo, Li Shen 等NeurIPS 2025 · 被引用 1 次
- MSeg3D: Multi-Modal 3D Semantic Segmentation for Autonomous DrivingJiale Li, Hang Dai, Hao Han, Yong DingCVPR 2023
- Seeing Through Fog Without Seeing Fog: Deep Multimodal Sensor Fusion in Unseen Adverse WeatherMario Bijelic, Tobias Gruber, Fahim Mannan, Florian Kraus 等CVPR 2020
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic SegmentationBowen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen 等NeurIPS 2025 · 被引用 8 次
- Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype DistillationJiaqi Tan, Xu Zheng, Yang LiuCVPR 2026 · 被引用 1 次
