ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
Haian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao, Jonathan T. Barron, Noah Snavely, Aleksander Holynski
Abstract
Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and have a computational cost that scales quadratically with the number of input images, making them inefficient when applied to large image collections. Sequential-reconstruction approaches reduce this cost but sacrifice reconstruction quality. We introduce ZipMap, a stateful feed-forward model that achieves linear-time, bidirectional 3D reconstruction while matching or surpassing the accuracy of quadratic-time methods. ZipMap employs test-time training layers to zip an entire image collection into a compact hidden scene state in a single forward pass, enabling reconstruction of over 700 frames in under 10 seconds on a single H100 GPU, more than faster than state-of-the-art methods such as VGGT. Moreover, we demonstrate the benefits of having a stateful representation in real-time scene-state querying and its extension to sequential streaming reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c79d6c5b-138d-44fb-a757-0bf5e2fe8123Cited by top-tier papers2
- Long-Tail Internet Photo ReconstructionYuan Li, Yuanbo Xiangli, Hadar Averbuch-Elor, Noah Snavely et al.CVPR 2026 · 4 citations
- VGGT-ΩJianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev et al.CVPR 2026
Builds on39
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
Related papers
- STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerYushi Lan, Yihang Luo, Fangzhou Hong, Shangchen Zhou et al.ICLR 2026 · 84 citations
- Block-Sparse Global Attention for Efficient Multi-View Geometry TransformersChung-Shien Brian Wang, Christian Schmidt, Jens Piekenbrinck, Bastian LeibeCVPR 2026 · 22 citations
- Streaming Visual Geometry TransformerDong Zhuo, Wenzhao Zheng, Jiahe Guo, Yuqi Wu et al.ICLR 2026 · 109 citations
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 14 citations
- VGG-T: Offline Feed-Forward 3D Reconstruction at ScaleSven Elflein, Ruilong Li, Sérgio Agostinho, Zan Gojcic et al.CVPR 2026
