DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
Chieh Hubert Lin, Zhaoyang Lv, Songyin Wu, Zhen Xu, Thu Nguyen-Phuoc, Hung-Yu Tseng, Julian Straub, Numair Khan, Lei Xiao, Ming-Hsuan Yang, Yuheng Ren, Richard A. Newcombe
Abstract
We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction has gained significant attention for its ability to rapidly create digital replicas of real-world environments. However, most existing models are limited to static scenes and fail to reconstruct the motion of moving objects. Developing a feed-forward model for dynamic scene reconstruction poses significant challenges, including the scarcity of training data and the need for appropriate 3D representations and training paradigms. To address these challenges, we introduce several key technical contributions: an enhanced large-scale synthetic dataset with ground-truth multi-view videos and dense 3D scene flow supervision; a per-pixel deformable 3D Gaussian representation that is easy to learn, supports high-quality dynamic view synthesis, and enables long-range 3D tracking; and a large transformer network that achieves real-time, generalizable dynamic scene reconstruction. Extensive qualitative and quantitative experiments demonstrate that DGS-LRM achieves dynamic scene reconstruction quality comparable to optimization-based methods, while significantly outperforming the state-of-the-art predictive dynamic reconstruction method on real-world examples. Its predicted physically grounded 3D deformation is accurate and can readily adapt for long-range 3D tracking tasks, achieving performance on par with state-of-the-art monocular video 3D tracking methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- PhysGM: Large Physical Gaussian Model for Feed-Forward 4D SynthesisChunji Lv, Zequn Chen, Donglin Di, Weinan Zhang et al.CVPR 2026 · 9 citations
- UFO-4D: Unposed Feedforward 4D Reconstruction from Two ImagesJunhwa Hur, Charles Herrmann, Songyou Peng, Philipp Henzler et al.ICLR 2026 · 5 citations
- Restage4D: Reanimating Deformable 3D Reconstruction from a Single VideoJixuan He, Chieh Hubert Lin, Lu Qi, Ming-Hsuan YangNeurIPS 2025 · 2 citations
- RetimeGS: Continuous-Time Reconstruction of 4D Gaussian SplattingXuezhen Wang, Li Ma, Yulin Shen, Zeyu Wang et al.CVPR 2026
- ReFlow: Self-correction Motion Learning for Dynamic Scene ReconstructionYanzhe Liang, Ruijie Zhu, Hanzhi Chang, Zhuoyuan Li et al.CVPR 2026
Builds on53
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
Related papers
- Motion Decoupled 3D Gaussian Splatting for Dynamic Object RepresentationXiao Hu, Libo Long, Jochen LangAAAI 2025 · 2 citations
- TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable TokensJiawei Ren, Michal J. Tyszkiewicz, Jiahui Huang, Zan GojcicCVPR 2026 · 13 citations
- StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video StreamsZike Wu, Qi Yan, Xuanyu Yi, Lele Wang et al.ICLR 2026 · 9 citations
- Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular VideosHanxue Liang, Jiawei Ren, Ashkan Mirzaei, Antonio Torralba et al.NeurIPS 2025 · 52 citations
- tttLRM: Test-Time Training for Long Context and Autoregressive 3D ReconstructionChen Wang, Hao Tan, Wang Yifan, Zhiqin Chen et al.CVPR 2026 · 11 citations
