RGB↔X: Image decomposition and synthesis using material- and lighting-aware diffusion models
Zheng Zeng, Valentin Deschaintre, Iliyan Georgiev, Yannick Hold-Geoffroy, Yiwei Hu, Fujun Luan, Ling-Qi Yan, Milos Hasan
Abstract
The three areas of realistic forward rendering, per-pixel inverse rendering, and generative image synthesis may seem like separate and unrelated sub-fields of graphics and vision. However, recent work has demonstrated improved estimation of per-pixel intrinsic channels (albedo, roughness, metallicity) based on a diffusion architecture; we call this the RGB → X problem. We further show that the reverse problem of synthesizing realistic images given intrinsic channels, X → RGB, can also be addressed in a diffusion framework. Focusing on the image domain of interior scenes, we introduce an improved diffusion model for RGB → X, which also estimates lighting, as well as the first diffusion X → RGB model capable of synthesizing realistic images from (full or partial) intrinsic channels. Our X → RGB model explores a middle ground between traditional rendering and generative models: We can specify only certain appearance properties that should be followed, and give freedom to the model to hallucinate a plausible version of the rest. This flexibility allows using a mix of heterogeneous training datasets that differ in the available channels. We use multiple existing datasets and extend them with our own synthetic and real data, resulting in a model capable of extracting scene properties better than previous work and of generating highly realistic images of interior scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers75
- UniRelight: Learning Joint Decomposition and Synthesis for Video RelightingKai He, Ruofan Liang, Jacob Munkberg, Jon Hasselgren et al.NeurIPS 2025 · 42 citations
- IntrinsiX: High-Quality PBR Generation using Image PriorsPeter Kocsis, Lukas Höllein, Matthias NießnerNeurIPS 2025 · 20 citations
- LuxDiT: Lighting Estimation with Video Diffusion TransformerRuofan Liang, Kai He, Zan Gojcic, Igor Gilitschenski et al.NeurIPS 2025 · 20 citations
- GaSLight: Gaussian Splats for Spatially-Varying Lighting in HDRChristophe Bolduc, Yannick Hold-Geoffroy, Jean-François LalondeICCV 2025 · 14 citations
- MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow GenerationKerui Ren, Jiayang Bai, Linning Xu, Lihan Jiang et al.NeurIPS 2025 · 9 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor ScenesJunyong Choi, Min-Cheol Sagong, SeokYeong Lee, Seung-Won Jung et al.CVPR 2025
- Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion ModelsRuofan Liang, Zan Gojcic, Huan Ling, Jacob Munkberg et al.CVPR 2025
- DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse RenderingRongjia Zheng, Qing Zhang, Chengjiang Long, Wei-Shi ZhengICCV 2025 · 2 citations
- Intrinsic Image Diffusion for Indoor Single-view Material EstimationPeter Kocsis, Vincent Sitzmann, Matthias NießnerCVPR 2024
- Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream DiffusionZhifei Chen, Tianshuo Xu, Wenhang Ge, Leyi Wu et al.CVPR 2025
