Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
Chuhan Zhang, Guillaume Le Moing, Skanda Koppula, Ignacio Rocco, Liliane Momeni, Junyu Xie, Shuyang Sun, Rahul Sukthankar, Joëlle K. Barral, Raia Hadsell, Zoubin Ghahramani, Andrew Zisserman
Abstract
Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently solve this task. D4RT utilizes a unified transformer architecture to jointly infer depth, spatiotemporal correspondence, and full camera parameters from a single video. Its core innovation is a novel querying mechanism that sidesteps the heavy computation of dense, perframe decoding and the complexity of managing multiple, task-specific decoders. Our decoding interface allows the model to independently and flexibly probe the 3D position of any point in space and time. The result is a lightweight and highly scalable method that enables remarkably efficient training and inference. We demonstrate that our approach sets a new state of the art, outperforming previous methods across a wide spectrum of 4D reconstruction tasks. We refer to the project webpage for animated results. 3
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8709ade9-ff9f-49b7-89ef-574eb36d4569Cited by top-tier papers6
- In Pursuit of Pixel Supervision for Visual Pre-trainingLihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei et al.CVPR 2026 · 13 citations
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan et al.ICML 2026 · 12 citations
- DynaTok: Token-Based 4D Reconstruction from Partial Point CloudsWeirong Chen, Keisuke Tateno, Hidenobu Matsuki, Michael Niemeyer et al.ICML 2026
- ReFlow: Self-correction Motion Learning for Dynamic Scene ReconstructionYanzhe Liang, Ruijie Zhu, Hanzhi Chang, Zhuoyuan Li et al.CVPR 2026
- VGGT-ΩJianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev et al.CVPR 2026
Builds on28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
Related papers
- Complet4R: Geometric Complete 4D ReconstructionWeibang Wang, Kenan Li, Zhuoguang Chen, Yijun Yuan et al.CVPR 2026
- C4D: 4D Made from 3D Through Dual CorrespondencesShizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao WangICCV 2025 · 4 citations
- PAGE-4D: Disentangled Pose and Geometry Estimation for VGGT-4D PerceptionKaichen Zhou, Yuhan Wang, Grace Chen, Gaspard Beaudouin et al.ICLR 2026 · 12 citations
- MoRe: Motion-aware Feed-forward 4D Reconstruction TransformerJuntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di et al.CVPR 2026 · 8 citations
- Motion 3-to-4: 3D Motion Reconstruction for 4D SynthesisHongyuan Chen, Xingyu Chen, Zexiang Xu, Anpei ChenCVPR 2026 · 17 citations
