Light of Normals: Unified Feature Representation for Universal Photometric Stereo
Houyuan Chen, Hong Li, Chongjie Ye, Zhaoxi Chen, Bohan Li, Shaocong Xu, Xianda Guo, Xuhui Liu, Yikai Wang, Baochang Zhang, Satoshi Ikehata, Boxin Shi
摘要
Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination models. Despite progress (e.g., SDM UniPS), two challenges remain. First, current encoders cannot guarantee that illumination and normal information are decoupled. To enforce decoupling, we introduce LINO UniPS with two key components: (i) Light Register Tokens with light alignment supervision to aggregate point, direction, and environment lights; (ii) Interleaved Attention Block featuring global cross-image attention that takes all lighting conditions together so the encoder can factor out lighting while retaining normal-related evidence. Second, high-frequency geometric details are easily lost. We address this with (i) a Wavelet-based Dual-branch Architecture and (ii) a Normal-gradient Perception Loss. These techniques yield a unified feature space in which lighting is explicitly represented by register tokens, while normal details are preserved via wavelet branch. We further introduce PS-Verse, a large-scale synthetic dataset graded by geometric complexity and lighting diversity, and adopt curriculum training from simple to complex scenes. Extensive experiments show new state-of-the-art results on public benchmarks (e.g., DiLiGenT, Luces), stronger generalization to real materials, and improved efficiency; ablations confirm that Light Register Tokens + Interleaved Attention Block drive better feature decoupling, while Wavelet-based Dual-branch Architecture + Normal-gradient Perception Loss recover finer details.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PAGE-4D: Disentangled Pose and Geometry Estimation for VGGT-4D PerceptionKaichen Zhou, Yuhan Wang, Grace Chen, Gaspard Beaudouin 等ICLR 2026 · 被引用 12 次
- NeAR: Coupled Neural Asset-Renderer StackHong Li, Chongjie Ye, Houyuan Chen, Weiqing Xiao 等CVPR 2026 · 被引用 4 次
- Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo Under Limited Multi-Illumination CuesKing-Man Tam, Satoshi Ikehata, Yuta Asano, Zhaoyi An 等AAAI 2026
- Relit-LiVE: Relight Video by Jointly Learning Environment VideoWeiqing Xiao, Hong Li, Xiuyu Yang, Houyuan Chen 等SIGGRAPH 2026
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu 等NeurIPS 2022 · 被引用 924 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt TrainingXiaoyang Wu, Zhuotao Tian, Xin Wen, Bohao Peng 等CVPR 2024 · 被引用 39 次
相关 Paper
- Universal Photometric Stereo Network using Global Lighting ContextsSatoshi IkehataCVPR 2022 · 被引用 22 次
- Spin-UP: Spin Light for Natural Light Uncalibrated Photometric StereoZongrui Li, Zhan Lu, Haojie Yan, Boxin Shi 等CVPR 2024
- Learning Efficient Photometric Feature Transform for Multi-view StereoKaizhang Kang, Cihui Xie, Ruisheng Zhu, Xiaohe Ma 等ICCV 2021 · 被引用 3 次
- DiLiGenT-Π: Photometric Stereo for Planar Surfaces with Rich Details - Benchmark Dataset and BeyondFeishi Wang, Jieji Ren, Heng Guo, Mingjun Ren 等ICCV 2023 · 被引用 18 次
- PX-NET: Simple and Efficient Pixel-Wise Training of Photometric Stereo NetworksFotios Logothetis, Ignas Budvytis, Roberto Mecca, Roberto CipollaICCV 2021 · 被引用 60 次
