POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction
Songyan Zhang, Yongtao Ge, Jinyuan Tian, Guangkai Xu, Hao Chen, Chen Lv, Chunhua Shen
摘要
Recent approaches to 3D reconstruction in dynamic scenes primarily rely on the integration of separate geometry estimation and matching modules, where the latter plays a critical role in distinguishing dynamic regions and mitigating the interference caused by moving objects. Furthermore, the matching module explicitly models object motion, enabling the tracking of specific targets and advancing motion understanding in complex scenarios. Recently, the proposed representation of pointmap in DUSt3R suggests a potential solution to unify both geometry estimation and matching in 3D space, effectively reducing computational overhead by eliminating the need for redundant auxiliary modules. However, it still struggles with ambiguous correspondences in dynamic regions, which limits reconstruction performance in such scenarios. In this work, we present POMATO, a unified framework for dynamic 3D reconstruction by marrying POintmap MAtching with Temporal mOtion. Specifically, our method first learns an explicit matching relationship by mapping RGB pixels across different views to 3D pointmaps within a unified coordinate system. Furthermore, we introduce a temporal motion module for dynamic motions that ensures scale consistency across different frames and enhances performance in 3D reconstruction tasks requiring both precise geometry and reliable matching, most notably 3D point tracking. We show the effectiveness of our proposed POMATO by demonstrating the remarkable performance across multiple downstream tasks, including video depth estimation, 3D point tracking, and pose estimation. Code and models are publicly available at https://github.com/wyddmw/POMATO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta 等CVPR 2026 · 被引用 35 次
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 被引用 17 次
- MoRe: Motion-aware Feed-forward 4D Reconstruction TransformerJuntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di 等CVPR 2026 · 被引用 8 次
- MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAERuijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han 等CVPR 2026 · 被引用 3 次
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular VideosShi Chen, Erik Sandström, Sandro Lombardi, Siyuan Li 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper24
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
相关 Paper
- Enhancing 3D Reconstruction for Dynamic ScenesJisang Han, Honggyu An, Jaewoo Jung, Takuya Narihira 等NeurIPS 2025 · 被引用 11 次
- C4D: 4D Made from 3D Through Dual CorrespondencesShizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao WangICCV 2025 · 被引用 4 次
- Dynamic Point Maps: A Versatile Representation for Dynamic 3D ReconstructionEdgar Sucar, Zihang Lai, Eldar Insafutdinov, Andrea VedaldiICCV 2025 · 被引用 9 次
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionJunyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani 等ICLR 2025 · 被引用 3 次
- V-DPM: 4D Video Reconstruction with Dynamic Point MapsEdgar Sucar, Eldar Insafutdinov, Zihang Lai, Andrea VedaldiCVPR 2026 · 被引用 29 次
