TTT3R: 3D Reconstruction as Test-Time Training
Xingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger, Anpei Chen
Abstract
have long been the foundation for 3D structure reconstruction and camera pose estimation. These methods rely on associating 2D correspondences [4, 25, 47, 50, 63] or minimizing reprojected photometric errors [29, 30] , followed by bundle adjustment (BA) [1, 12, 77, 79, 80, 86] for structure and motion refinement. Although highly effective when assembled into comprehensive systems [50, 66] , these approaches often struggle in conditions of small camera parallax or ill-posed conditions (e.g., dynamic or textureless), leading to performance degradation. Recent work, such as MegaSaM [43] and VIPE [35] , has demonstrated progress in adapting traditional SLAM paradigms to dynamic scenes by integrating semantic segmentation [35, 39] , optical flows [35, 43, 104, 105] , and geometric constraints [35, 39, 43, 48, 104] . Concurrently, methods like VGGT-SLAM [49] and seek improved robustness by integrating learned front-ends [51, 87, 91] . However, these methods require iterative optimization based on off-the-shelf estimation, where synchronization barriers often lead to cumulative errors and high computational overhead. This reliance hinders real-time online inference and learning scalability (e.g., the 'tabula rasa' blank slate limitation [89] ). In this work, we investigate data-driven feed-forward models with generalizable priors to enable dense 3D reconstruction even from dynamic and textureless video sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e76133bf-5038-4d9c-b6b1-5bfe43e1d281Cited by top-tier papers31
- Human3R: Everyone Everywhere All at OnceYue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen et al.ICLR 2026 · 38 citations
- Imagine360: Immersive 360 Video Generation from Perspective AnchorJing Tan, Shuai Yang, Tong Wu, Jingwen He et al.NeurIPS 2025 · 33 citations
- Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionTao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai et al.CVPR 2026 · 26 citations
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time TrainingHaian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao et al.CVPR 2026 · 23 citations
- Motion 3-to-4: 3D Motion Reconstruction for 4D SynthesisHongyuan Chen, Xingyu Chen, Zexiang Xu, Anpei ChenCVPR 2026 · 17 citations
Builds on45
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- WildPose: A Unified Framework for Robust Pose Estimation in the WildJianhao Zheng, Liyuan Zhu, Zihan Zhu, Iro ArmeniCVPR 2026 · 1 citation
- Dynamic Visual SLAM using a General 3D PriorXingguang Zhong, Liren Jin, Marija Popovic, Jens Behley et al.CVPR 2026 · 1 citation
- MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic VideosZhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang et al.CVPR 2025
- Back on Track: Bundle Adjustment for Dynamic Scene ReconstructionWeirong Chen, Ganlin Zhang, Felix Wimbauer, Rui Wang et al.ICCV 2025 · 1 citation
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular VideosShi Chen, Erik Sandström, Sandro Lombardi, Siyuan Li et al.NeurIPS 2025 · 1 citation
