IRISformer: Dense Vision Transformers for Single-Image Inverse Rendering in Indoor Scenes
Rui Zhu, Zhengqin Li, Janarbek Matai, Fatih Porikli, Manmohan Chandraker
Abstract
Indoor scenes exhibit significant appearance variations due to myriad interactions between arbitrarily diverse object shapes, spatially-changing materials, and complex lighting. Shadows, highlights, and inter-reflections caused by visible and invisible light sources require reasoning about long-range interactions for inverse rendering, which seeks to recover the components of image formation, namely, shape, material, and lighting. In this work, our intuition is that the long-range attention learned by transformer architectures is ideally suited to solve longstanding challenges in single-image inverse rendering. We demonstrate with a specific instantiation of a dense vision transformer, , that excels at both single-task and multi-task reasoning required for inverse rendering. Specifically, we propose a transformer architecture to simultaneously estimate depths, normals, spatially-varying albedo, roughness and lighting from a single image of an indoor scene. Our extensive evaluations on benchmark datasets demonstrate state-of-the-art results on each of the above tasks, enabling applications like object insertion and material editing in a single unconstrained real image, with greater photorealism than prior works. Code and data are publicly released at https://github.com/ViLab-UCSD/IRISformer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d9ee3d9-eb4b-4f97-ba9d-e4b205997b52Cited by top-tier papers30
- IntrinsicNeRF: Learning Intrinsic Neural Radiance Fields for Editable Novel View SynthesisWeicai Ye, Shuo Chen, Chong Bao, Hujun Bao et al.ICCV 2023 · 60 citations
- EverLight: Indoor-Outdoor Editable HDR Lighting EstimationMohammad Reza Karimi Dastjerdi, Jonathan Eisenmann, Yannick Hold-Geoffroy, Jean-François LalondeICCV 2023 · 41 citations
- Factorized Inverse Path Tracing for Efficient and Accurate Material-Lighting EstimationLiwen Wu, Rui Zhu, Mustafa B. Yaldiz, Yinhao Zhu et al.ICCV 2023 · 27 citations
- MatSynth: A Modern PBR Materials DatasetGiuseppe Vecchio, Valentin DeschaintreCVPR 2024 · 24 citations
- Intrinsic Image Fusion for Multi-View 3D Material ReconstructionPeter Kocsis, Lukas Höllein, Matthias NießnerCVPR 2026 · 8 citations
Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Neural Inverse Rendering of an Indoor Scene From a Single ImageSoumyadip Sengupta, Jinwei Gu, Kihwan Kim, Guilin Liu et al.ICCV 2019 · 172 citations
- Learning Indoor Inverse Rendering with 3D Spatially-Varying LightingZian Wang, Jonah Philion, Sanja Fidler, Jan KautzICCV 2021 · 109 citations
- IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range ImagesChih-Hao Lin, Jia-Bin Huang, Zhengqin Li, Zhao Dong et al.CVPR 2025
- Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF From a Single ImageZhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli et al.CVPR 2020
- MAIR: Multi-View Attention Inverse Rendering with 3D Spatially-Varying Lighting EstimationJunyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung et al.CVPR 2023
