Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation
Jiawei Han, Matteo Poggi, Li Huan, Changshuo Wang, Kaiqi Liu, Wei Li
Abstract
The effective integration and utilization of multimodal data acquired from image cameras and LiDAR is of paramount importance for perception systems. This paper proposes I mage-to- P oint Cloud F eature Back- P rojection ( IPFP ), a novel method for training multimodal fusion networks that back-projects aggregated image-feature centers (from non-projection-aligned image pixels) into the point-cloud feature set via the estimated depth map. Consequently, image features and point cloud features reside within the same three-dimensional space, enabling the natural enrichment of image information into the point cloud during the network forward pass. This process can be selectively enabled when desired -- for instance, at training time -- and turned off in the absence of multimodal data -- for example, at testing time if only LiDAR sensors are available. Experimental results demonstrate that IPFP can consistently improve state-of-the-art 3D semantic segmentation models, while retaining the ability to process LiDAR-only data at inference time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on37
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- Learning Multi-View Aggregation In the Wild for Large-Scale 3D Semantic SegmentationDamien Robert, Bruno Vallet, Loïc LandrieuCVPR 2022 · 84 citations
- MFINet: Multi-view Fusion and 2D-3D Interaction Enhancement for Real-Time LiDAR Semantic SegmentationNan Ma, Zhijie Liu, Yiheng HanAAAI 2026
- DeepI2P: Image-to-Point Cloud Registration via Deep ClassificationJiaxin Li, Gim Hee LeeCVPR 2021
- Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic SegmentationZhuangwei Zhuang, Rong Li, Kui Jia, Qicheng Wang et al.ICCV 2021 · 129 citations
- PointPainting: Sequential Fusion for 3D Object DetectionSourabh Vora, Alex H. Lang, Bassam Helou, Oscar BeijbomCVPR 2020
