LiTo: Surface Light Field Tokenization
Jen-Hao Rick Chang, Xiaoming Zhao, Dorian Chan, Oncel Tuzel
Abstract
We propose a 3D latent representation that jointly models object geometry and view-dependent appearance. Most prior works focus on either reconstructing 3D geometry or predicting view-independent diffuse appearance, and thus struggle to capture realistic view-dependent effects. Our approach leverages that RGB-depth images provide samples of a surface light field. By encoding random subsamples of this surface light field into a compact set of latent vectors, our model learns to represent both geometry and appearance within a unified 3D latent space. This representation reproduces view-dependent effects such as specular highlights and Fresnel reflections under complex lighting. We further train a latent flow matching model on this representation to learn its distribution conditioned on a single input image, enabling the generation of 3D objects with appearances consistent with the lighting and materials in the input. Experiments show that our approach achieves higher visual quality and better input fidelity than existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch et al.ICLR 2022 · 797 citations
Related papers
- Orchid: Image Latent Diffusion for Joint Appearance and Geometry GenerationAkshay Krishnan, Xinchen Yan, Vincent Casser, Abhijit KunduICCV 2025 · 8 citations
- SViM3D: Stable Video Material Diffusion for Single Image 3D GenerationAndreas Engelhardt, Mark Boss, Vikram Voleti, Chun-Han Yao et al.ICCV 2025 · 1 citation
- SpecNeRF: Gaussian Directional Encoding for Specular ReflectionsLi Ma, Vasu Agrawal, Haithem Turki, Changil Kim et al.CVPR 2024 · 13 citations
- Learning Neural Light Fields with Ray-Space EmbeddingBenjamin Attal, Jia-Bin Huang, Michael Zollhöfer, Johannes Kopf et al.CVPR 2022 · 75 citations
- Gen3R: 3D Scene Generation Meets Feed-Forward ReconstructionJiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma et al.CVPR 2026 · 24 citations
