Mask3D: Pretraining 2D Vision Transformers by Learning Masked 3D Priors
Ji Hou, Xiaoliang Dai, Zijian He, Angela Dai, Matthias Nießner
Abstract
Figure 1. We present Mask3D, which learns to embed 3D priors to 2D representations for image understanding tasks, based on a selfsupervised pre-training formulation from single RGB-D views, without requiring any camera pose or multi-view correspondence information. Our pre-training takes masked RGB and depth patches as input to reconstruct the dense depth map, and the pre-trained color backbone is used to fine-tune various downstream image understanding tasks. This results in effective ViT pre-training for a variety of downstream tasks and datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41dd4863-d40e-4290-9954-ba08fe9f15ddCited by top-tier papers7
- UnScene3D: Unsupervised 3D Instance Segmentation for Indoor ScenesDávid Rozenberszki, Or Litany, Angela DaiCVPR 2024 · 25 citations
- Multi-View Representation is What You Need for Point-Cloud Pre-TrainingSiming Yan, Chen Song, Youkang Kong, Qixing HuangICLR 2024 · 6 citations
- Unsupervised Semantic Segmentation Through Depth-Guided Feature Correlation and SamplingLeon Sick, Dominik Engel, Pedro Hermosilla, Timo RopinskiCVPR 2024 · 6 citations
- A Mixed Diet Makes DINO An Omnivorous Vision EncoderRishabh Kabra, Maks Ovsjanikov, Drew A. Hudson, Ye Xia et al.CVPR 2026 · 3 citations
- CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular FusionShoubin Yu, Jaehong Yoon, Mohit BansalICLR 2025
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
Related papers
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- Cross-View and Cross-Pose Completion for 3D Human UnderstandingMatthieu Armando, Salma Galaaoui, Fabien Baradel, Thomas Lucas et al.CVPR 2024
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
- GLID: Pre-training a Generalist Encoder-Decoder Vision ModelJihao Liu, Jinliang Zheng, Yu Liu, Hongsheng LiCVPR 2024 · 5 citations
- Semantically-Guided Representation Learning for Self-Supervised Monocular DepthVitor Guizilini, Rui Hou, Jie Li, Rares Ambrus et al.ICLR 2020 · 264 citations
