Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-Motion
Guoyu Lu
Abstract
Scene reconstruction has a wide range of applications in computer vision and robotics. To build practical constraints and feature correspondences, rich textures and distinguished gradient variations are particularly required in classic and learning-based SfM. When building lowtexture regions with repeated patterns, especially mostlywhite indoor rooms, there is a significant drop in performance. In this work, we propose Shading-SfM-Net, a novel framework for simultaneously learning a shape-fromshading network based on the inverse rendering constraint and a structure-from-motion framework based on warped keypoint, room layout, and geometric consistency, to improve structure-from-motion and surface reconstruction for low-texture indoor scenes. Shading-SfM-Net tightly incorporates the surface shape consistency and 3D geometric registration loss in order to dig into their mutual information and further overcome the instability on flat regions. We evaluate the proposed framework on texture-less indoor scenes (NYUv2 and ScanNet), and show that by simultaneously learning shading, motion and shape, our pipeline is able to achieve state-of-the-art performance with superior generalization capability for unseen texture-less datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 130273d4-3277-4eda-a184-e8a56449aa89Cited by top-tier papers2
- Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without TrainingHexiao Lu, Xiaokun Sun, Zeyu Cai, Hao Guo et al.CVPR 2026
- Vision-Language Embodiment for Monocular Depth EstimationJinchang Zhang, Guoyu LuCVPR 2025
Builds on21
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- HR-Depth: High Resolution Self-Supervised Monocular Depth EstimationXiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong et al.AAAI 2021 · 341 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
Related papers
- Detector-Free Structure from MotionXingyi He, Jiaming Sun, Yifan Wang, Sida Peng et al.CVPR 2024
- Geo-Neus: Geometry-Consistent Neural Implicit Surfaces Learning for Multi-view ReconstructionQiancheng Fu, Qingshan Xu, Yew Soon Ong, Wenbing TaoNeurIPS 2022 · 336 citations
- Joint Texture and Geometry Optimization for RGB-D ReconstructionYanping Fu, Qingan Yan, Jie Liao, Chunxia XiaoCVPR 2020
- Shape Anchor Guided Holistic Indoor Scene UnderstandingMingyue Dong, Linxi Huan, Hanjiang Xiong, Shuhan Shen et al.ICCV 2023 · 5 citations
- IM360: Large-Scale Indoor Mapping with 360 CamerasDongki Jung, Jaehoon Choi, Yonghan Lee, Dinesh ManochaICCV 2025 · 3 citations
