TrajVG: 3D Trajectory-Coupled Visual Geometry Learning
Xingyu Miao, Weiguang Zhao, Tao Lu, Linning Xu, Mulin Yu, Yang Long, Jiangmiao Pang, Junting Dong
摘要
Feed-forward multi-frame 3D reconstruction models often degrade on videos with object motion. Global-reference becomes ambiguous under multiple motions, while the local pointmap relies heavily on estimated relative poses and can drift, causing cross-frame misalignment and duplicated structures. We propose TrajVG, a reconstruction framework that makes cross-frame 3D correspondence an explicit prediction by estimating camera-coordinate 3D trajectories. We couple sparse trajectories, per-frame local point maps, and relative camera poses with geometric consistency objectives: (i) bidirectional trajectory–pointmap consistency with controlled gradient flow, and (ii) a pose consistency objective driven by static track anchors that suppresses gradients from dynamic regions. To scale training to in-the-wild videos where 3D trajectory labels are scarce, we reformulate the same coupling constraints into self-supervised objectives using only pseudo 2D tracks, enabling unified training with mixed supervision. Extensive experiments across 3D tracking, pose estimation, point-map reconstruction, and video depth show that TrajVG is particularly effective in challenging video settings with motion, weak overlap, or pose ambiguity, while remaining competitive on standard feed-forward reconstruction benchmarks. Project page: https://xingy038.github.io/TrajVG/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper36
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone 等ICCV 2021 · 被引用 686 次
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 被引用 659 次
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang 等ICLR 2026 · 被引用 318 次
相关 Paper
- Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild VideosMatthew Strong, Wei-Jer Chang, Quentin Herau, Jiezhi Yang 等CVPR 2026 · 被引用 3 次
- CamGeo: Sparse Camera-Conditioned Image-to-Video Generation with 3D Geometry PriorsXuanyi Liu, Deyi Ji, Liqun Liu, Lanyun Zhu 等ICML 2026 · 被引用 3 次
- Flow3r: Factored Flow Prediction for Scalable Visual Geometry LearningZhongxiao Cong, Qitao Zhao, Minsik Jeon, Shubham TulsianiCVPR 2026 · 被引用 8 次
- GGPT: Geometry-Grounded Point TransformerYutong Chen, Yiming Wang, Xucong Zhang, Sergey Prokudin 等CVPR 2026 · 被引用 2 次
- Track3R: Joint Point Map and Trajectory Prior for Spatiotemporal 3D UnderstandingSeong Hyeon Park, Jinwoo ShinNeurIPS 2025
