Multiview Compressive Coding for 3D Reconstruction
Chao-Yuan Wu, Justin Johnson, Jitendra Malik, Christoph Feichtenhofer, Georgia Gkioxari
Abstract
A central goal of visual recognition is to understand objects and scenes from a single image. 2D recognition has witnessed tremendous progress thanks to large-scale learning and general-purpose representations. Comparatively, 3D poses new challenges stemming from occlusions not depicted in the image. Prior works try to overcome these by inferring from multiple views or rely on scarce CAD models and category-specific priors which hinder scaling to novel settings. In this work, we explore single-view 3D reconstruction by learning generalizable representations inspired by advances in self-supervised learning. We introduce a simple framework that operates on 3D points of single objects or whole scenes coupled with category-agnostic largescale training from diverse RGB-D videos. Our model, Multiview Compressive Coding (MCC), learns to compress the input appearance and geometry to predict the 3D structure by querying a 3D-aware decoder. MCC's generality and efficiency allow it to learn from large-scale and diverse data sources with strong generalization to novel objects imagined by DALL•E 2 or captured in-the-wild with an iPhone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b275c9c1-d3d8-4802-845e-538950192192Cited by top-tier papers40
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai et al.ICLR 2024 · 973 citations
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi et al.ICLR 2024 · 813 citations
- One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape OptimizationMinghua Liu, Chao Xu, Haian Jin, Linghao Chen et al.NeurIPS 2023 · 755 citations
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Pretrain, Self-train, Distill: A simple recipe for Supersizing 3D ReconstructionKalyan Vasudev Alwala, Abhinav Gupta, Shubham TulsianiCVPR 2022 · 23 citations
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- Discovering 3D Parts from Image CollectionsChun-Han Yao, Wei-Chih Hung, Varun Jampani, Ming-Hsuan YangICCV 2021 · 21 citations
- ShapeClipper: Scalable 3D Shape Learning from Single-View Images via Geometric and CLIP-Based ConsistencyZixuan Huang, Varun Jampani, Anh Thai, Yuanzhen Li et al.CVPR 2023
- Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature SpaceLeonhard Sommer, Olaf Dünkel, Christian Theobalt, Adam KortylewskiCVPR 2025
