NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
Yanrui Bin, Wenbo Hu, Haoyuan Wang, Xinya Chen, Bing Wang
Abstract
Surface normal estimation serves as a cornerstone for a spectrum of computer vision applications. While numerous efforts have been devoted to static image scenarios, ensuring temporal coherence in video-based normal estimation remains a formidable challenge. Instead of merely augmenting existing methods with temporal components, we present NormalCrafter to leverage the inherent temporal priors of Video Diffusion Models (VDMs). We identify the reason for blurry predictions when directly applying VDMs and introduce Semantic Feature Regularization (SFR) to encourage the model to concentrate on geometric details by aligning diffusion features with fine-grained semantic cues. Moreover, we introduce a two-stage training protocol that leverages both latent and pixel space learning to preserve spatial accuracy while maintaining long temporal context. Extensive evaluations demonstrate the efficacy of our method, showcasing a superior performance in generating temporally consistent normal sequences with intricate details from diverse videos. Code and models are publicly available at https://normalcrafter.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Monocular Normal Estimation via Shading Sequence EstimationZongrui Li, Xinhua Ma, Minghui Hu, Yunqing Zhao et al.ICLR 2026 · 3 citations
- Generative Perception of Shape and Material from Differential MotionXinran Nicole Han, Ko Nishino, Todd E. ZicklerNeurIPS 2025 · 2 citations
- ReflFlow: Learning Geometry-Guided Ray Tracing for Dynamic Specular ReconstructionJiachen Tao, Junyi Wu, Haoxuan Wang, Zongxin Yang et al.ICML 2026
- Relit-LiVE: Relight Video by Jointly Learning Environment VideoWeiqing Xiao, Hong Li, Xiuyu Yang, Houyuan Chen et al.SIGGRAPH 2026
Builds on18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsSongchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie et al.ICCV 2025 · 6 citations
- Geometrycrafter: Consistent Geometry Estimation for Open-World Videos With Diffusion PriorsTian-Xing Xu, Xiangjun Gao, Wenbo Hu, Xiaoyu Li et al.ICCV 2025 · 3 citations
- GeoVideo: Introducing Geometric Regularization into Video Generation ModelYunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang et al.NeurIPS 2025 · 18 citations
- MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAERuijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han et al.CVPR 2026 · 3 citations
- DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationYuejiang Dong, Wang Zhao, Jiale Xu, Ying Shan et al.ICCV 2025
