Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation
Su Sun, Cheng Zhao, Himangi Mittal, Gaurav Mittal, Rohith Kukkala, Yingjie Victor Chen, Mei Chen
摘要
Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize that the view discrepancy arises from supervision limited to pixel- or latent-space video-diffusion losses, which lack explicitly temporally aware, feature-level tracking guidance. We present Track4DGen, a two-stage framework that couples a multi-view video diffusion model with a foundation point tracker and a hybrid 4D Gaussian Splatting (4D-GS) reconstructor. The central idea is to explicitly inject tracker-derived motion priors into intermediate feature representations for both multi-view video generation and 4D-GS. In Stage One, we enforce dense, feature-level point correspondences inside the diffusion generator, producing temporally consistent features that curb appearance drift and enhance cross-view coherence. In Stage Two, we reconstruct a dynamic 4D-GS using a hybrid motion encoding that concatenates co-located diffusion features (carrying Stage-One tracking priors) with Hex-plane features, and augment them with 4D Spherical Harmonics for higher-fidelity dynamics modeling. Track4DGen surpasses baselines on both multi-view video generation and 4D generation benchmarks, yielding temporally stable, text-editable 4D assets. Lastly, we curate Sketchfab28, a high-quality dataset for benchmarking object-centric 4D generation and fostering future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai 等ICLR 2024 · 被引用 973 次
- Representation Alignment for Diffusion Transformers without External ComponentsDengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang 等ICLR 2026 · 被引用 532 次
相关 Paper
- Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video GenerationHyeonho Jeong, Chun-Hao P. Huang, Jong Chul Ye, Niloy J. Mitra 等CVPR 2025
- EG4D: Explicit Generation of 4D Object without Score DistillationQi Sun, Zhiyang Guo, Ziyu Wan, Jing Nathan Yan 等ICLR 2025
- Gaussian Variation Field Diffusion for High-Fidelity Video-to-4D SynthesisBowen Zhang, Sicheng Xu, Chuxin Wang, Jiaolong Yang 等ICCV 2025 · 被引用 4 次
- Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion ModelsHanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang 等NeurIPS 2024 · 被引用 116 次
- Sparse4DGS: Flow-Geometry Assisted 4D Gaussian Splatting for Dynamic Sparse View SynthesisDongdong Hu, Yang Zhou, Xiaofeng Huang, Haibing Yin 等ACM MM 2025 · 被引用 3 次
