ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models
Lukas Höllein, Aljaz Bozic, Norman Müller, David Novotný, Hung-Yu Tseng, Christian Richardt, Michael Zollhöfer, Matthias Nießner
2024Year
38Top-tier citations
Abstract
2 Meta https://lukashoel.github.io/ViewDiff/ a stuffed bear sitting on a wooden box Input Multi-view generated images Figure 1. Multi-view consistent image generation. Our method takes as input a text description, or any number of posed input images, and generates high-quality, multi-view consistent images of a real-world 3D object in authentic surroundings from any desired camera poses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers38
- CAT3D: Create Anything in 3D with Multi-View Diffusion ModelsRuiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee et al.NeurIPS 2024 · 490 citations
- Meta 3D AssetGen: Text-to-Mesh Generation with High-Quality Geometry, Texture, and PBR MaterialsYawar Siddiqui, Tom Monnier, Filippos Kokkinos, Mahendra Kariya et al.NeurIPS 2024 · 89 citations
- MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D EditingChenjie Cao, Chaohui Yu, Fan Wang, Xiangyang Xue et al.NeurIPS 2024 · 36 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- IntrinsiX: High-Quality PBR Generation using Image PriorsPeter Kocsis, Lukas Höllein, Matthias NießnerNeurIPS 2025 · 20 citations
Builds on44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Consistent View Synthesis with Pose-Guided Diffusion ModelsHung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan et al.CVPR 2023
- ReplaceAnything3D: Text-Guided Object Replacement in 3D Scenes with Compositional Scene RepresentationsEdward Bartrum, Thu Nguyen-Phuoc, Christopher Xie, Zhengqin Li et al.NeurIPS 2024 · 5 citations
- ConsistNet: Enforcing 3D Consistency for Multi-View Images DiffusionJiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji et al.CVPR 2024 · 25 citations
- MIRROR: Make Your Object-Level Multi-View Generation More Consistent with Training-Free RectificationTianchi Xing, Bonan Li, Congying Han, Xinmin Qiu et al.ICML 2025
- Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar CreationXiyi Chen, Marko Mihajlovic, Shaofei Wang, Sergey Prokudin et al.CVPR 2024
