Learning Multi-View Aggregation In the Wild for Large-Scale 3D Semantic Segmentation
Damien Robert, Bruno Vallet, Loïc Landrieu
Abstract
Recent works on 3D semantic segmentation propose to exploit the synergy between images and point clouds by processing each modality with a dedicated network and projecting learned 2D features onto 3D points. Merging large-scale point clouds and images raises several challenges, such as constructing a mapping between points and pixels, and aggregating features between multiple views. Current methods require mesh reconstruction or specialized sensors to recover occlusions, and use heuristics to select and aggregate available images. In contrast, we propose an end-to-end trainable multi-view aggregation model leveraging the viewing conditions of 3D points to merge features from images taken at arbitrary positions. Our method can combine standard 2D and 3D networks and outperforms both 3D models operating on colorized point clouds and hybrid 2D/3D networks without requiring colorization, meshing, or true depth maps. We set a new state-of-the-art for large-scale indoor/outdoor semantic segmentation on S3DIS (74.7 mIoU 6-Fold) and on KITTI-360 (58.3 mIoU). Our full pipeline is accessible at https: //github.com/drprojects/DeepViewAgg, and only requires raw 3D scans and a set of images and poses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1185f69e-92b1-45af-aac8-66cd9a2d51e1Cited by top-tier papers21
- Efficient 3D Semantic Segmentation with Superpoint TransformerDamien Robert, Hugo Raguet, Loïc LandrieuICCV 2023 · 131 citations
- HUGS: Holistic Urban 3D Scene Understanding via Gaussian SplattingHongyu Zhou, Jiahao Shao, Lu Xu, Dongfeng Bai et al.CVPR 2024 · 50 citations
- 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level SupervisionCheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang, Yen-Yu LinICCV 2023 · 30 citations
- MaskClustering: View Consensus Based Mask Graph Clustering for Open-Vocabulary 3D Instance SegmentationMi Yan, Jiazhao Zhang, Yan Zhu, He WangCVPR 2024 · 22 citations
- Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceJinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 18 citations
Builds on9
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- MVTN: Multi-View Transformation Network for 3D Shape RecognitionAbdullah Hamdi, Silvio Giancola, Bernard GhanemICCV 2021 · 280 citations
- TransformerFusion: Monocular RGB Scene Reconstruction using TransformersAljaz Bozic, Pablo R. Palafox, Justus Thies, Angela Dai et al.NeurIPS 2021 · 185 citations
- Learning Relationships for Multi-View 3D Object RecognitionZe Yang, Liwei WangICCV 2019 · 166 citations
Related papers
- ODIN: A Single Model for 2D and 3D SegmentationAyush Jain, Pushkal Katara, Nikolaos Gkanatsios, Adam W. Harley et al.CVPR 2024
- Robust Multi-Modality Multi-Object TrackingWenwei Zhang, Hui Zhou, Shuyang Sun, Zhe Wang et al.ICCV 2019 · 221 citations
- Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic SegmentationXuweiyi Chen, Wentao Zhou, Aruni RoyChowdhury, Zezhou ChengICLR 2026 · 4 citations
- Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic SegmentationJiawei Han, Matteo Poggi, Li Huan, Changshuo Wang et al.CVPR 2026
- Multi-View Representation is What You Need for Point-Cloud Pre-TrainingSiming Yan, Chen Song, Youkang Kong, Qixing HuangICLR 2024 · 6 citations
