MWSIS: Multimodal Weakly Supervised Instance Segmentation with 2D Box Annotations for Autonomous Driving
Guangfeng Jiang, Jun Liu, Yuzhi Wu, Wenlong Liao, Tao He, Pai Peng
Abstract
Instance segmentation is a fundamental research in computer vision, especially in autonomous driving. However, manual mask annotation for instance segmentation is quite time-consuming and costly. To address this problem, some prior works attempt to apply weakly supervised manner by exploring 2D or 3D boxes. However, no one has ever successfully segmented 2D and 3D instances simultaneously by only using 2D box annotations, which could further reduce the annotation cost by an order of magnitude. Thus, we propose a novel framework called Multimodal Weakly Supervised Instance Segmentation (MWSIS), which incorporates various fine-grained label correction modules for both 2D and 3D modalities, along with a new multimodal cross-supervision approach. In the 2D pseudo label generation branch, the Instance-based Pseudo Mask Generation (IPG) module utilizes predictions for self-supervised correction. Similarly, in the 3D pseudo label generation branch, the Spatial-based Pseudo Label Generation (SPG) module generates pseudo labels by incorporating the spatial prior information of the point cloud. To further refine the generated pseudo labels, the Point-based Voting Label Correction (PVC) module utilizes historical predictions for correction. Additionally, a Ring Segment-based Label Correction (RSC) module is proposed to refine the predictions by leveraging the depth prior information from the point cloud. Finally, the Consistency Sparse Cross-modal Supervision (CSCS) module reduces the inconsistency of multimodal predictions by response distillation. Particularly, transferring the 3D backbone to downstream tasks not only improves the performance of the 3D detectors, but also outperforms fully supervised instance segmentation with only 5% fully supervised annotations. On the Waymo dataset, the proposed framework demonstrates significant improvements over the baseline, especially achieving 2.59% mAP and 12.75% mAP increases for 2D and 3D instance segmentation tasks, respectively. The code is available at https://github.com/jiangxb98/mwsis-plugin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e720215-1418-4fa8-b35f-c7ae386a6db4Cited by top-tier papers3
- TRACE: Your Diffusion Model is Secretly an Instance Edge DetectorSanghyun Jo, Ziseok Lee, Wooyeol Lee, Jonghyun Choi et al.ICLR 2026 · 4 citations
- PromptMoE: A Segmentation Refinement Framework Leveraging Mixture of Experts for Improved PromptingStephen Price, Danielle L. Cote, Elke A. RundensteinerCVPR 2026
- ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance SegmentationShengchao Zhou, Jiehong Lin, Jiahui Liu, Shizhen Zhao et al.AAAI 2026
Builds on20
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Fully Sparse 3D Object DetectionLue Fan, Feng Wang, Naiyan Wang, Zhaoxiang ZhangNeurIPS 2022 · 168 citations
- Pointly-Supervised Instance SegmentationBowen Cheng, Omkar Parkhi, Alexander KirillovCVPR 2022 · 140 citations
- Scribble-Supervised LiDAR Semantic SegmentationOzan Unal, Dengxin Dai, Luc Van GoolCVPR 2022 · 86 citations
Related papers
- LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous DrivingXiang Li, Junbo Yin, Botian Shi, Yikang Li et al.AAAI 2023 · 16 citations
- Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial LabelsHaizhuang Liu, Huimin Ma, Yilin Wang, Bochao Zou et al.ACM MM 2022 · 6 citations
- Sketchy Bounding-box Supervision for 3D Instance SegmentationQian Deng, Le Hui, Jin Xie, Jian YangCVPR 2025
- SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic PromptsShijia Zhao, Qiming Xia, Xusheng Guo, Pufan Zou et al.CVPR 2025
- Seg2Box: 3D Object Detection by Point-Wise Semantics SupervisionMaoji Zheng, Ziyu Xu, Qiming Xia, Hai Wu et al.AAAI 2025 · 3 citations
