3D-to-2D Distillation for Indoor Scene Parsing
Zhengzhe Liu, Xiaojuan Qi, Chi-Wing Fu
Abstract
Indoor scene semantic parsing from RGB images is very challenging due to occlusions, object distortion, and viewpoint variations. Going beyond prior works that leverage geometry information, typically paired depth maps, we present a new approach, a 3D-to-2D distillation framework, that enables us to leverage 3D features extracted from largescale 3D data repositories (e.g., ScanNet-v2) to enhance 2D features extracted from RGB images. Our work has three novel contributions. First, we distill 3D knowledge from a pretrained 3D network to supervise a 2D network to learn simulated 3D features from 2D features during the training, so the 2D network can infer without requiring 3D data. Second, we design a two-stage dimension normalization scheme to calibrate the 2D and 3D features for better integration. Third, we design a semantic-aware adversarial training model to extend our framework for training with unpaired 3D data. Extensive experiments on various datasets, ScanNet-V2, S3DIS, and NYU-v2, demonstrate the superiority of our approach. Also, experimental results show that our 3D-to-2D distillation improves the model generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Towards Efficient 3D Object Detection with Knowledge DistillationJihan Yang, Shaoshuai Shi, Runyu Ding, Zhe Wang et al.NeurIPS 2022 · 76 citations
- Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisXu Yan, Heshen Zhan, Chaoda Zheng, Jiantao Gao et al.NeurIPS 2022 · 49 citations
- EgoDistill: Egocentric Head Motion Distillation for Efficient Video UnderstandingShuhan Tan, Tushar Nagarajan, Kristen GraumanNeurIPS 2023 · 44 citations
- ImLoveNet: Misaligned Image-supported Registration Network for Low-overlap Point Cloud PairsHonghua Chen, Zeyong Wei, Yabin Xu, Mingqiang Wei et al.SIGGRAPH 2022 · 29 citations
- Transferring CLIP's Knowledge into Zero-Shot Point Cloud Semantic SegmentationYuanbin Wang, Shaofei Huang, Yulu Gao, Zhen Wang et al.ACM MM 2023 · 17 citations
Builds on3
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- OccuSeg: Occupancy-Aware 3D Instance SegmentationLei Han, Tian Zheng, Lan Xu, Lu FangCVPR 2020
- PointGroup: Dual-Set Point Grouping for 3D Instance SegmentationLi Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu et al.CVPR 2020
Related papers
- 3D Segmenter: 3D Transformer based Semantic Segmentation via 2D Panoramic DistillationZhennan Wu, Yang Li, Yifei Huang, Lin Gu et al.ICLR 2023
- Pri3D: Can 3D Priors Help 2D Representation Learning?Ji Hou, Saining Xie, Benjamin Graham, Angela Dai et al.ICCV 2021 · 94 citations
- Geometry-Aware Network for Domain Adaptive Semantic SegmentationYinghong Liao, Wending Zhou, Xu Yan, Zhen Li et al.AAAI 2023 · 9 citations
- GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic SegmentationXujing Tao, Chuxin Wang, Yubo Ai, Zhixin Cheng et al.CVPR 2026 · 3 citations
- PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel ClusteringZisheng Chen, Hongbin Xu, Weitao Chen, Zhipeng Zhou et al.ICCV 2023 · 21 citations
