Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis
Hongyuan Chen, Xingyu Chen, Zexiang Xu, Anpei Chen
Abstract
We present Motion 3-to-4, a feed-forward framework for synthesising high-quality 4D dynamic objects from a single monocular video and an optional 3D reference mesh. While recent advances have significantly improved 2D, video, and 3D content generation, 4D synthesis remains difficult due to limited training data and the inherent ambiguity of recovering geometry and motion from a monocular viewpoint. Motion 3-to-4 addresses these challenges by decomposing 4D synthesis into static 3D shape generation and motion reconstruction. Using a canonical reference mesh, our model learns a compact motion latent representation and predicts per-frame vertex trajectories to recover complete, temporally coherent geometry. A scalable frame-wise transformer further enables robustness to varying sequence lengths. Evaluations on both standard benchmarks and a new dataset with accurate ground-truth geometry show that Motion 3-to-4 delivers superior fidelity and spatial consis-tency compared to prior work. Project page is available at https://motion3-to-4.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d9695a8-5d98-4102-bdf2-a6949d3e4999Cited by top-tier papers1
Ask how each one uses itBuilds on74
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
Related papers
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan et al.ICML 2026 · 12 citations
- Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular VideoZeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus et al.CVPR 2026 · 13 citations
- Inferring Compositional 4D Scenes without Ever Seeing OneAhmet Berke Gökmen, Ajad Chhatkuli, Luc Van Gool, Danda PaudelCVPR 2026 · 1 citation
- ShapeGen4D: Towards High Quality 4D Shape Generation from VideosJiraphon Yenphraphai, Ashkan Mirzaei, Jianqi Chen, Jiaxu Zou et al.ICLR 2026 · 21 citations
- MoRe: Motion-aware Feed-forward 4D Reconstruction TransformerJuntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di et al.CVPR 2026 · 8 citations
