MWSIS: Multimodal Weakly Supervised Instance Segmentation with 2D Box Annotations for Autonomous Driving
Guangfeng Jiang, Jun Liu, Yuzhi Wu, Wenlong Liao, Tao He, Pai Peng
摘要
Instance segmentation is a fundamental research in computer vision, especially in autonomous driving. However, manual mask annotation for instance segmentation is quite time-consuming and costly. To address this problem, some prior works attempt to apply weakly supervised manner by exploring 2D or 3D boxes. However, no one has ever successfully segmented 2D and 3D instances simultaneously by only using 2D box annotations, which could further reduce the annotation cost by an order of magnitude. Thus, we propose a novel framework called Multimodal Weakly Supervised Instance Segmentation (MWSIS), which incorporates various fine-grained label correction modules for both 2D and 3D modalities, along with a new multimodal cross-supervision approach. In the 2D pseudo label generation branch, the Instance-based Pseudo Mask Generation (IPG) module utilizes predictions for self-supervised correction. Similarly, in the 3D pseudo label generation branch, the Spatial-based Pseudo Label Generation (SPG) module generates pseudo labels by incorporating the spatial prior information of the point cloud. To further refine the generated pseudo labels, the Point-based Voting Label Correction (PVC) module utilizes historical predictions for correction. Additionally, a Ring Segment-based Label Correction (RSC) module is proposed to refine the predictions by leveraging the depth prior information from the point cloud. Finally, the Consistency Sparse Cross-modal Supervision (CSCS) module reduces the inconsistency of multimodal predictions by response distillation. Particularly, transferring the 3D backbone to downstream tasks not only improves the performance of the 3D detectors, but also outperforms fully supervised instance segmentation with only 5% fully supervised annotations. On the Waymo dataset, the proposed framework demonstrates significant improvements over the baseline, especially achieving 2.59% mAP and 12.75% mAP increases for 2D and 3D instance segmentation tasks, respectively. The code is available at https://github.com/jiangxb98/mwsis-plugin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TRACE: Your Diffusion Model is Secretly an Instance Edge DetectorSanghyun Jo, Ziseok Lee, Wooyeol Lee, Jonghyun Choi 等ICLR 2026 · 被引用 4 次
- PromptMoE: A Segmentation Refinement Framework Leveraging Mixture of Experts for Improved PromptingStephen Price, Danielle L. Cote, Elke A. RundensteinerCVPR 2026
- ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance SegmentationShengchao Zhou, Jiehong Lin, Jiahui Liu, Shizhen Zhao 等AAAI 2026
它引用的顶会 Paper20
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Fully Sparse 3D Object DetectionLue Fan, Feng Wang, Naiyan Wang, Zhaoxiang ZhangNeurIPS 2022 · 被引用 168 次
- Pointly-Supervised Instance SegmentationBowen Cheng, Omkar Parkhi, Alexander KirillovCVPR 2022 · 被引用 140 次
- Scribble-Supervised LiDAR Semantic SegmentationOzan Unal, Dengxin Dai, Luc Van GoolCVPR 2022 · 被引用 86 次
相关 Paper
- LWSIS: LiDAR-Guided Weakly Supervised Instance Segmentation for Autonomous DrivingXiang Li, Junbo Yin, Botian Shi, Yikang Li 等AAAI 2023 · 被引用 16 次
- Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial LabelsHaizhuang Liu, Huimin Ma, Yilin Wang, Bochao Zou 等ACM MM 2022 · 被引用 6 次
- Sketchy Bounding-box Supervision for 3D Instance SegmentationQian Deng, Le Hui, Jin Xie, Jian YangCVPR 2025
- SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic PromptsShijia Zhao, Qiming Xia, Xusheng Guo, Pufan Zou 等CVPR 2025
- Seg2Box: 3D Object Detection by Point-Wise Semantics SupervisionMaoji Zheng, Ziyu Xu, Qiming Xia, Hai Wu 等AAAI 2025 · 被引用 3 次
