Bidirectional Projection Network for Cross Dimension Scene Understanding
Wenbo Hu, Hengshuang Zhao, Li Jiang, Jiaya Jia, Tien-Tsin Wong
Abstract
2D image representations are in regular grids and can be processed efficiently, whereas 3D point clouds are unordered and scattered in 3D space. The information inside these two visual domains is well complementary, e.g., 2D images have fine-grained texture while 3D point clouds contain plentiful geometry information. However, most current visual recognition systems process them individually. In this paper, we present a bidirectional projection network (BPNet) for joint 2D and 3D reasoning in an end-to-end manner. It contains 2D and 3D sub-networks with symmetric architectures, that are connected by our proposed bidirectional projection module (BPM). Via the BPM, complementary 2D and 3D information can interact with each other in multiple architectural levels, such that advantages in these two visual domains can be combined for better scene recognition. Extensive quantitative and qualitative experimental evaluations show that joint reasoning over 2D and 3D visual domains can benefit both 2D and 3D scene understanding simultaneously. Our BPNet achieves top performance on the ScanNetV2 benchmark for both 2D and 3D semantic segmentation. Code is available at https://github.com/wbhu/BPNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers36
- Neural 3D Scene Reconstruction with the Manhattan-world AssumptionHaoyu Guo, Sida Peng, Haotong Lin, Qianqian Wang et al.CVPR 2022 · 152 citations
- Pri3D: Can 3D Priors Help 2D Representation Learning?Ji Hou, Saining Xie, Benjamin Graham, Angela Dai et al.ICCV 2021 · 94 citations
- Learning Multi-View Aggregation In the Wild for Large-Scale 3D Semantic SegmentationDamien Robert, Bruno Vallet, Loïc LandrieuCVPR 2022 · 84 citations
- VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic SegmentationZeyu Hu, Xuyang Bai, Jiaxiang Shang, Runze Zhang et al.ICCV 2021 · 78 citations
- X -Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense CaptioningZhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo et al.CVPR 2022 · 72 citations
Builds on12
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Hierarchical Point-Edge Interaction Network for Point Cloud Semantic SegmentationLi Jiang, Hengshuang Zhao, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 213 citations
- xMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic SegmentationMaximilian Jaritz, Tuan-Hung Vu, Raoul de Charette, Émilie Wirbel et al.CVPR 2020
Related papers
- DSPNet: Dual-vision Scene Perception for Robust 3D Question AnsweringJingzhou Luo, Yang Liu, Weixing Chen, Zhen Li et al.CVPR 2025
- PointMBF: A Multi-scale Bidirectional Fusion Network for Unsupervised RGB-D Point Cloud RegistrationMingzhi Yuan, Kexue Fu, Zhihao Li, Yucong Meng et al.ICCV 2023 · 29 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Joint Learning of 2D-3D Weakly Supervised Semantic SegmentationHyeokjun Kweon, Kuk-Jin YoonNeurIPS 2022 · 34 citations
- PointGroup: Dual-Set Point Grouping for 3D Instance SegmentationLi Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu et al.CVPR 2020
