ORBIT: Benchmarking SfM in the Wild with 360° Video
Sara Sabour, Richard Tucker, Marcus Brubaker, Saurabh Saxena, Junhwa Hur, Andrea Tagliasacchi, Deqing Sun, David J. Fleet, Richard Szeliski, Noah Snavely
摘要
Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes.Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard to gauge real-world progress, or pinpoint where improvements are most needed.To address this gap, we introduce a new benchmark for evaluating camera pose estimation.Our key insight is to leverage online panoramic 360° as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery.The panoramic nature of these videos provides richer visual context for tracking camera motion, even when parts of the view are affected by blur, motion, or dynamic objects.By tracking camera motion across full 360° videos, we crop and reproject selected portions to generate perspective-view clips that serve as our benchmark---ORBIT---a diverse collection of 100 video clips.Experiments show that COLMAP and other state-of-the-art SfM methods struggle to accurately estimate camera positions on our benchmark, indicating that it remains a challenging and open problem space for future research.As a result, ORBIT provides a valuable testbed where researchers can meaningfully compete and measure progress on truly challenging, real-world SfM problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- Monocular Dynamic View Synthesis: A Reality CheckHang Gao, Ruilong Li, Shubham Tulsiani, Bryan Russell 等NeurIPS 2022 · 被引用 235 次
- SpatialVID: A Large-Scale Video Dataset with Spatial AnnotationsJiahao Wang, Yufeng Yuan, Rujie Zheng, Youtian Lin 等CVPR 2026 · 被引用 72 次
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos 等ICLR 2025 · 被引用 15 次
相关 Paper
- Robust Dynamic Radiance FieldsYu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng 等CVPR 2023
- Uncalibrated Structure from Motion on a SphereJonathan Ventura, Viktor Larsson, Fredrik KahlICCV 2025 · 被引用 1 次
- CasualStereo: Casual Capture of Stereo Panoramas with Spherical Structure-from-MotionLewis Baker, Steven Mills, Stefanie Zollmann, Jonathan VenturaIEEE VR 2020 · 被引用 2 次
- Global Structure-from-Motion Meets Feedforward ReconstructionLinfei Pan, Johannes L. Schönberger, Marc PollefeysCVPR 2026 · 被引用 5 次
- PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesDiantao Tu, Hainan Cui, Xianwei Zheng, Shuhan ShenCVPR 2024 · 被引用 4 次
