SceneTok: A Compressed, Diffusable Token Space for 3D Scenes
Mohammad Asim, Christopher Wewer, Jan Lenssen
Abstract
We present SceneTok, a novel tokenizer for encoding view sets of scenes into a compressed and diffusable set of unstructured tokens. Existing approaches for 3D scene representation and generation commonly use 3D data structures or view-aligned fields. In contrast, we introduce the first method that encodes scene information into a small set of permutation invariant tokens that is disentangled from the spatial grid. The scene tokens are predicted by a multi-view tokenizer given many context views and rendered into novel views by employing a light-weight rectified flow decoder. A diffusion transformer enables scene generation on the compressed token space. We show that the compression is two orders of magnitude stronger than for other representations while still reaching state-of-the-art reconstruction quality. Further, our representation can be rendered from novel trajectories, including ones deviating from the input trajectory, and we show that the decoder gracefully handles uncertainty. Finally, the highly-compressed set of unstructured latent scene tokens enables simple and efficient scene generation in 8 seconds, achieving a much better quality-speed tradeoff than previous paradigms.Our code and trained models will be released upon acceptance of the paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61cf641e-6a10-4acf-9fe0-289e5b178937Builds on49
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
Related papers
- RecTok: Reconstruction Distillation along Rectified FlowQingyu Shi, Size Wu, Jinbin Bai, Kaidong Yu et al.CVPR 2026 · 5 citations
- CLiFT: Compressive Light-Field Tokens for Compute Efficient and Adaptive Neural RenderingZhengqing Wang, Yuefan Wu, Jiacheng Chen, Fuyang Zhang et al.NeurIPS 2025 · 2 citations
- Scaling Transformer-Based Novel View Synthesis with Models Token Disentanglement and Synthetic DataNithin Gopalakrishnan Nair, Srinivas Kaza, Xuan Luo, Vishal M. Patel et al.ICCV 2025 · 1 citation
- UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D ReconstructionZhenwei Zhu, Liying Yang, Ning Li, Chaohao Jiang et al.ICCV 2023 · 12 citations
- Representing 3D Shapes with 64 Latent Vectors for 3D Diffusion ModelsIn Cho, Youngbeom Yoo, Subin Jeon, Seon Joo KimICCV 2025 · 1 citation
