SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
Andreas Engelhardt, Mark Boss, Vikram Voleti, Chun-Han Yao, Hendrik P. A. Lensch, Varun Jampani
摘要
We present Stable Video Materials 3D (SViM3D), a framework to predict multi-view consistent physically based rendering (PBR) materials, given a single image. Recently, video diffusion models have been successfully used to reconstruct 3D objects from a single image efficiently. However, reflectance is still represented by simple material models or needs to be estimated in additional steps to enable relighting and controlled appearance edits. We extend a latent video diffusion model to output spatially varying PBR parameters and surface normals jointly with each generated view based on explicit camera control. This unique setup allows for relighting and generating a 3D asset using our model as neural prior. We introduce various mechanisms to this pipeline that improve quality in this ill-posed setting. We show state-of-the-art relighting and novel view synthesis performance on multiple object-centric datasets. Our method generalizes to diverse inputs, enabling the generation of relightable 3D assets useful in AR/VR, movies, games and other visual media.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- NeAR: Coupled Neural Asset-Renderer StackHong Li, Chongjie Ye, Houyuan Chen, Weiqing Xiao 等CVPR 2026 · 被引用 4 次
- MatSpray: Fusing 2D Material World Knowledge on 3D GeometryPhilipp Langsteiner, Jan-Niklas Dihlmann, Hendrik P. A. LenschCVPR 2026 · 被引用 2 次
- Property-Informed Diffusion-Based Text-to-Microstructure GenerationBingxuan Dai, Hongsong Wang, Jie GuiCVPR 2026
它引用的顶会 Paper51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- GR3EN: Generative Relighting for 3D EnvironmentsXiaoyan Xing, Philipp Henzler, Junhwa Hur, Runze Li 等SIGGRAPH 2026
- Relit-LiVE: Relight Video by Jointly Learning Environment VideoWeiqing Xiao, Hong Li, Xiuyu Yang, Houyuan Chen 等SIGGRAPH 2026
- Single Image Neural Material RelightingJames C. Bieron, Xin Tong, Pieter PeersSIGGRAPH 2023 · 被引用 3 次
- ReLi3D: Relightable Multi-view 3D Reconstruction with Disentangled IlluminationJan-Niklas Dihlmann, Mark Boss, Simon Donné, Andreas Engelhardt 等ICLR 2026 · 被引用 1 次
- GAS: Generative Avatar Synthesis from a Single ImageYixing Lu, Junting Dong, Youngjoong Kwon, Qin Zhao 等ICCV 2025 · 被引用 5 次
