Multimodal Material Segmentation
Yupeng Liang, Ryosuke Wakaki, Shohei Nobuhara, Ko Nishino
Abstract
Recognition of materials from their visual appearance is essential for computer vision tasks, especially those that involve interaction with the real world. Material segmentation, i.e., dense per-pixel recognition of materials, remains challenging as, unlike objects, materials do not exhibit clearly discernible visual signatures in their regular RGB appearances. Different materials, however, do lead to different radiometric behaviors, which can often be captured with non-RGB imaging modalities. We realize multimodal material segmentation from RGB, polarization, and near-infrared images. We introduce the MCubeS dataset (from MultiModal Material Segmentation) which contains 500 sets of multimodal images capturing 42 street scenes. Ground truth material segmentation as well as semantic segmentation are annotated for every image and pixel. We also derive a novel deep neural network, MCubeSNet, which learns to focus on the most informative combinations of imaging modalities for each material class with a newly derived region-guided filter selection (RGFS) layer. We use semantic segmentation as a prior to “guide” this filter selection. To the best of our knowledge, our work is the first comprehensive study on truly multimodal material segmentation. We believe our work opens new avenues of practical use of material information in safety critical applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4664a305-ed79-41df-8f19-38e00bb24bfeCited by top-tier papers25
- StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic SegmentationBingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao et al.ACM MM 2025 · 15 citations
- Hierarchical Material Recognition from Local AppearanceMatthew Beveridge, Shree K. NayarICCV 2025 · 5 citations
- Polarization Guided Mask-Free Shadow RemovalChu Zhou, Chao Xu, Boxin ShiAAAI 2025 · 4 citations
- MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge DistillationHui Li, Pengfei Yang, Juanyang Chen, Le Dong et al.ACM MM 2025 · 4 citations
- MMCert: Provable Defense Against Adversarial Attacks to Multi-Modal ModelsYanting Wang, Hongye Fu, Wei Zou, Jinyuan JiaCVPR 2024 · 4 citations
Builds on5
- Surface Normals and Shape From WaterSatoshi Murai, Meng-Yu Kuo, Ryo Kawahara, Shohei Nobuhara et al.ICCV 2019 · 14 citations
- Dynamic Region-Aware ConvolutionJin Chen, Xijun Wang, Zichao Guo, Xiangyu Zhang et al.CVPR 2021
- Multi-Modal Fusion Transformer for End-to-End Autonomous DrivingAditya Prakash, Kashyap Chitta, Andreas GeigerCVPR 2021
- Decoupled Dynamic Filter NetworksJingkai Zhou, Varun Jampani, Zhixiong Pi, Qiong Liu et al.CVPR 2021
- MMTM: Multimodal Transfer Module for CNN FusionHamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, Kazuhito KoishidaCVPR 2020
Related papers
- Glass Segmentation using Intensity and Spectral Polarization CuesHaiyang Mei, Bo Dong, Wen Dong, Jiaxi Yang et al.CVPR 2022 · 91 citations
- Exploiting Polarized Material Cues for Robust Car DetectionWen Dong, Haiyang Mei, Ziqi Wei, Ao Jin et al.AAAI 2024 · 10 citations
- Multi-Sensor Large-Scale Dataset for Multi-View 3D ReconstructionOleg Voynov, Gleb Bobrovskikh, Pavel A. Karpyshev, Saveliy Galochkin et al.CVPR 2023
- PolarDepth: Monocular Transparent Object Depth from Polar-Physics PriorsWen Dong, Haiyang Mei, Yinglian Ji, Zijun Zhang et al.ICML 2026
- SemanticRT: A Large-Scale Dataset and Method for Robust Semantic Segmentation in Multispectral ImagesWei Ji, Jingjing Li, Cheng Bian, Zhicheng Zhang et al.ACM MM 2023 · 22 citations
