4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos
Zhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou, Richard A. Newcombe, Zhaoyang Lv
Abstract
Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1) Gaussian 2 (t-1)
Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 2 (t+1) Gaussian 1 Gaussian 3 Veloci ty Velocity Rotation Velocit y t life-span t1 DinoV2 Dynamic Gaussian Reconstruction Figure 1: We propose a scalable 4D dynamic reconstruction model trained only on real-world monocular RGB videos. The feed-forward 4DGS (section 3.1) representation enables us to render the geometry and appearance of the dynamic scene from novel views in real-time. Even without explicit supervision, the model can learn to distinguish dynamic contents from the background and produce realistic optical flows. The figure shows an enlarged set of Gaussians for the purpose of visualization. The embedded rendered videos only play in Adobe Reader or KDE Okular.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ea3051f-9199-4cf0-b8e1-6f00b32d3c43Cited by top-tier papers11
- NeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosYuxue Yang, Lue Fan, Ziqi Shi, Junran Peng et al.CVPR 2026 · 42 citations
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng et al.ICCV 2025 · 29 citations
- DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed ImagesXiaoxue Chen, Ziyi Xiong, Yuantao Chen, Gen Li et al.CVPR 2026 · 24 citations
- ActionMesh: Animated 3D Mesh Generation with Temporal 3D DiffusionRemy Sabathier, David Novotný, Niloy J. Mitra, Tom MonnierCVPR 2026 · 16 citations
- ART: Articulated Reconstruction TransformerZizhang Li, Cheng Zhang, Zhengqin Li, Henry Howard-Jenkins et al.CVPR 2026 · 12 citations
Builds on33
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi et al.ICLR 2024 · 813 citations
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
Related papers
- DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular VideosChieh Hubert Lin, Zhaoyang Lv, Songyin Wu, Zhen Xu et al.NeurIPS 2025 · 15 citations
- Gaussian-Flow: 4D Reconstruction with Dynamic 3D Gaussian ParticleYoutian Lin, Zuozhuo Dai, Siyu Zhu, Yao YaoCVPR 2024
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- Flux4D: Flow-based Unsupervised 4D ReconstructionJingkang Wang, Henry Che, Yun Chen, Ze Yang et al.NeurIPS 2025 · 10 citations
- Deblur4DGS: 4D Gaussian Splatting from Blurry Monocular VideoRenlong Wu, Zhilu Zhang, Mingyang Chen, Zifei Yan et al.AAAI 2026 · 17 citations
