Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
Mukund Varma T., Peihao Wang, Zhiwen Fan, Zhangyang Wang, Hao Su, Ravi Ramamoorthi
摘要
In recent years, there has been an explosion of 2D vision models for numerous tasks such as semantic segmentation, style transfer or scene editing, enabled by large-scale 2D image datasets. At the same time, there has been renewed interest in 3D scene representations such as neural radiance fields from multi-view images. However, the availability of 3D or multiview data is still substantially limited compared to 2D image datasets, making extending 2D vision models to 3D data highly desirable but also very challenging. Indeed, extending a single 2D vision operator like scene editing to 3D typically requires a highly creative method specialized to that task and often requires per-scene optimization. In this paper, we ask the question of whether any 2D vision model can be lifted to make 3D consistent predictions. We answer this question in the affirmative; our new Lift3D method trains to predict unseen views onfeature spaces generated by afew visual models (i.e. DINO and CLIP), but then generalizes to novel vision operators and tasks, such as style transfer, super-resolution, open vocabulary segmentation and image colorization; for some of these tasks, there is no comparable previous 3D method. In many cases, we even outperform state-of-the-art methods specialized for the task in question. Moreover, Lift3D is a zero-shot method, in the sense that it requires no task-specific training, nor scene-specific optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Large Spatial Model: End-to-end Unposed Images to Semantic 3DZhiwen Fan, Jian Zhang, Wenyan Cong, Peihao Wang 等NeurIPS 2024 · 被引用 86 次
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial RepresentationsYujia Zhang, Xiaoyang Wu, Yixing Lao, Chengyao Wang 等NeurIPS 2025 · 被引用 47 次
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva 等ICCV 2025 · 被引用 4 次
- Splattalk: 3D VQA with Gaussian SplattingAnh Thai, Songyou Peng, Kyle Genova, Leonidas J. Guibas 等ICCV 2025 · 被引用 4 次
- LatentVLA: Taming Latent Space for Generalizable and Long-Horizon Bimanual ManipulationJunming WangAAAI 2026 · 被引用 1 次
它引用的顶会 Paper32
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- Weakly Supervised 3D Open-vocabulary SegmentationKunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu 等NeurIPS 2023 · 被引用 173 次
- Decomposing NeRF for Editing via Feature Field DistillationSosuke Kobayashi, Eiichi Matsumoto, Vincent SitzmannNeurIPS 2022 · 被引用 479 次
- FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation ModelsJianglong Ye, Naiyan Wang, Xiaolong WangICCV 2023 · 被引用 56 次
- Putting NeRF on a Diet: Semantically Consistent Few-Shot View SynthesisAjay Jain, Matthew Tancik, Pieter AbbeelICCV 2021 · 被引用 615 次
- From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMsAng Cao, Sergio Arnaud, Oleksandr Maksymets, Jianing Yang 等ICML 2025
