Pri3D: Can 3D Priors Help 2D Representation Learning?
Ji Hou, Saining Xie, Benjamin Graham, Angela Dai, Matthias Nießner
Abstract
Recent advances in 3D perception have shown impressive progress in understanding geometric structures of 3D shapes and even scenes. Inspired by these advances in geometric understanding, we aim to imbue image-based perception with representations learned under geometric constraints. We introduce an approach to learn view-invariant, geometry-aware representations for network pre-training, based on multi-view RGB-D data, that can then be effectively transferred to downstream 2D tasks. We propose to employ contrastive learning under both multi-view image constraints and image-geometry constraints to encode 3D priors into learned 2D representations. This results not only in improvement over 2D-only representation learning on the image-based tasks of semantic segmentation, instance segmentation and object detection on real-world indoor datasets, but moreover, provides significant improvement in the low data regime. We show significant improvement of 6.0% on semantic segmentation on full data as well as 11.9% on 20% data against baselines on ScanNet. Our code is open sourced at https://github.com/Sekunde/Pri3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d2f0e56-4cbb-471d-8d52-3db596866192Cited by top-tier papers27
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- Segment Any Point Cloud Sequences by Distilling Vision Foundation ModelsYouquan Liu, Lingdong Kong, Jun Cen, Runnan Chen et al.NeurIPS 2023 · 169 citations
- Image-to-Lidar Self-Supervised Distillation for Autonomous Driving DataCorentin Sautier, Gilles Puy, Spyros Gidaris, Alexandre Boulch et al.CVPR 2022 · 102 citations
- Ponder: Point Cloud Pre-training via Neural RenderingDi Huang, Sida Peng, Tong He, Honghui Yang et al.ICCV 2023 · 55 citations
- Joint Learning of 2D-3D Weakly Supervised Semantic SegmentationHyeokjun Kweon, Kuk-Jin YoonNeurIPS 2022 · 34 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
Related papers
- Self-Supervised Image Representation Learning with Geometric Set ConsistencyNenglun Chen, Lei Chu, Hao Pan, Yan Lu et al.CVPR 2022 · 8 citations
- Multi-View Representation is What You Need for Point-Cloud Pre-TrainingSiming Yan, Chen Song, Youkang Kong, Qixing HuangICLR 2024 · 6 citations
- SimIPU: Simple 2D Image and 3D Point Cloud Unsupervised Pre-training for Spatial-Aware Visual RepresentationsZhenyu Li, Zehui Chen, Ang Li, Liangji Fang et al.AAAI 2022 · 78 citations
- Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastGuofan Fan, Zekun Qi, Wenkai Shi, Kaisheng MaACM MM 2024 · 12 citations
- DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationBowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu et al.ICLR 2024 · 110 citations
