2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision
Cheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang, Yen-Yu Lin
摘要
We present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3D features are complementary for point cloud segmentation. However, existing methods require extra 2D annotations to achieve 2D-3D information fusion. Considering the high annotation cost of point clouds, effective 2D and 3D feature fusion based on weakly supervised learning is in great demand. To this end, we propose a transformer model with two encoders and one decoder for weakly supervised point cloud segmentation using only scene-level class tags. Specifically, the two encoders compute the selfattended features for 3D point clouds and 2D multi-view images, respectively. The decoder implements interlaced 2D-3D cross-attention and carries out implicit 2D and 3D feature fusion. We alternately switch the roles of queries and key-value pairs in the decoder layers. It turns out that the 2D and 3D features are iteratively enriched by each other. Experiments show that it performs favorably against existing weakly supervised point cloud segmentation methods by a large margin on the S3DIS and Scan-Net benchmarks. The project page will be available at https://jimmy15923.github.io/mit_web/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd 等CVPR 2022 · 被引用 275 次
- Learning Relationships for Multi-View 3D Object RecognitionZe Yang, Liwei WangICCV 2019 · 被引用 166 次
相关 Paper
- An MIL-Derived Transformer for Weakly Supervised Point Cloud SegmentationCheng-Kun Yang, Ji-Jia Wu, Kai-Syun Chen, Yung-Yu Chuang 等CVPR 2022 · 被引用 53 次
- Joint Learning of 2D-3D Weakly Supervised Semantic SegmentationHyeokjun Kweon, Kuk-Jin YoonNeurIPS 2022 · 被引用 34 次
- Point Cloud Self-Supervised Learning via 3D to Multi-View Masked LearnerZhimin Chen, Xuewei Chen, Xiao Guo, Yingwei Li 等ICCV 2025 · 被引用 1 次
- Multi-Modality Affinity Inference for Weakly Supervised 3D Semantic SegmentationXiawei Li, Qingyuan Xu, Jing Zhang, Tianyi Zhang 等AAAI 2024 · 被引用 7 次
- SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D FeaturesJinyuan Qu, Hongyang Li, Xingyu Chen, Shilong Liu 等AAAI 2026 · 被引用 2 次
