SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
Andreas Engelhardt, Mark Boss, Vikram Voleti, Chun-Han Yao, Hendrik P. A. Lensch, Varun Jampani
Abstract
We present Stable Video Materials 3D (SViM3D), a framework to predict multi-view consistent physically based rendering (PBR) materials, given a single image. Recently, video diffusion models have been successfully used to reconstruct 3D objects from a single image efficiently. However, reflectance is still represented by simple material models or needs to be estimated in additional steps to enable relighting and controlled appearance edits. We extend a latent video diffusion model to output spatially varying PBR parameters and surface normals jointly with each generated view based on explicit camera control. This unique setup allows for relighting and generating a 3D asset using our model as neural prior. We introduce various mechanisms to this pipeline that improve quality in this ill-posed setting. We show state-of-the-art relighting and novel view synthesis performance on multiple object-centric datasets. Our method generalizes to diverse inputs, enabling the generation of relightable 3D assets useful in AR/VR, movies, games and other visual media.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbcef167-c5ee-4be3-8c5f-aa8c374c2afbCited by top-tier papers3
- NeAR: Coupled Neural Asset-Renderer StackHong Li, Chongjie Ye, Houyuan Chen, Weiqing Xiao et al.CVPR 2026 · 4 citations
- MatSpray: Fusing 2D Material World Knowledge on 3D GeometryPhilipp Langsteiner, Jan-Niklas Dihlmann, Hendrik P. A. LenschCVPR 2026 · 2 citations
- Property-Informed Diffusion-Based Text-to-Microstructure GenerationBingxuan Dai, Hongsong Wang, Jie GuiCVPR 2026
Builds on51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- GR3EN: Generative Relighting for 3D EnvironmentsXiaoyan Xing, Philipp Henzler, Junhwa Hur, Runze Li et al.SIGGRAPH 2026
- Relit-LiVE: Relight Video by Jointly Learning Environment VideoWeiqing Xiao, Hong Li, Xiuyu Yang, Houyuan Chen et al.SIGGRAPH 2026
- Single Image Neural Material RelightingJames C. Bieron, Xin Tong, Pieter PeersSIGGRAPH 2023 · 3 citations
- ReLi3D: Relightable Multi-view 3D Reconstruction with Disentangled IlluminationJan-Niklas Dihlmann, Mark Boss, Simon Donné, Andreas Engelhardt et al.ICLR 2026 · 1 citation
- GAS: Generative Avatar Synthesis from a Single ImageYixing Lu, Junting Dong, Youngjoong Kwon, Qin Zhao et al.ICCV 2025 · 5 citations
