UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
Junhwa Hur, Charles Herrmann, Songyou Peng, Philipp Henzler, Zeyu Ma, Todd Zickler, Deqing Sun
摘要
Dense 4D reconstruction from unposed images remains a critical challenge, with current methods relying on slow test-time optimization or fragmented, task-specific feedforward models. We introduce UFO-4D, a unified feedforward framework to reconstruct a dense, explicit 4D representation from just a pair of unposed images. UFO-4D directly estimates dynamic 3D Gaussian Splats, enabling the joint and consistent estimation of 3D geometry, 3D motion, and camera pose in a feedforward manner. Our core insight is that differentiably rendering multiple signals from a single Dynamic 3D Gaussian representation offers major training advantages. This approach enables a self-supervised image synthesis loss while tightly coupling appearance, depth, and motion. Since all modalities share the same geometric primitives, supervising one inherently regularizes and improves the others. This synergy overcomes data scarcity, allowing UFO-4D to outperform prior work by up to 3 times in joint geometry, motion, and camera pose estimation. Our representation also enables high-fidelity 4D interpolation across novel views and time. Please visit our project page for visual results: https://ufo-4d.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper36
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
相关 Paper
- Flux4D: Flow-based Unsupervised 4D ReconstructionJingkang Wang, Henry Che, Yun Chen, Ze Yang 等NeurIPS 2025 · 被引用 10 次
- DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous DrivingZhuolin He, Jing Li, Guanghao Li, Xiaolei Chen 等CVPR 2026 · 被引用 5 次
- Uncertainty Matters in Dynamic Gaussian Splatting for Monocular 4D ReconstructionFengzhi Guo, Chih-Chuan Hsu, Sihao Ding, Cheng ZhangICLR 2026 · 被引用 6 次
- Motion Decoupled 3D Gaussian Splatting for Dynamic Object RepresentationXiao Hu, Libo Long, Jochen LangAAAI 2025 · 被引用 2 次
- UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene ModelingKaiyuan Tan, Yingying Shen, Ziyue Zhu, Mingfei Tu 等CVPR 2026 · 被引用 4 次
