ORBIT: Benchmarking SfM in the Wild with 360° Video
Sara Sabour, Richard Tucker, Marcus Brubaker, Saurabh Saxena, Junhwa Hur, Andrea Tagliasacchi, Deqing Sun, David J. Fleet, Richard Szeliski, Noah Snavely
Abstract
Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes.Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard to gauge real-world progress, or pinpoint where improvements are most needed.To address this gap, we introduce a new benchmark for evaluating camera pose estimation.Our key insight is to leverage online panoramic 360° as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery.The panoramic nature of these videos provides richer visual context for tracking camera motion, even when parts of the view are affected by blur, motion, or dynamic objects.By tracking camera motion across full 360° videos, we crop and reproject selected portions to generate perspective-view clips that serve as our benchmark---ORBIT---a diverse collection of 100 video clips.Experiments show that COLMAP and other state-of-the-art SfM methods struggle to accurately estimate camera positions on our benchmark, indicating that it remains a challenging and open problem space for future research.As a result, ORBIT provides a valuable testbed where researchers can meaningfully compete and measure progress on truly challenging, real-world SfM problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be798716-647b-4bdd-9da7-b646017b3799Builds on14
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- Monocular Dynamic View Synthesis: A Reality CheckHang Gao, Ruilong Li, Shubham Tulsiani, Bryan Russell et al.NeurIPS 2022 · 235 citations
- SpatialVID: A Large-Scale Video Dataset with Spatial AnnotationsJiahao Wang, Yufeng Yuan, Rujie Zheng, Youtian Lin et al.CVPR 2026 · 72 citations
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos et al.ICLR 2025 · 15 citations
Related papers
- Robust Dynamic Radiance FieldsYu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng et al.CVPR 2023
- Uncalibrated Structure from Motion on a SphereJonathan Ventura, Viktor Larsson, Fredrik KahlICCV 2025 · 1 citation
- CasualStereo: Casual Capture of Stereo Panoramas with Spherical Structure-from-MotionLewis Baker, Steven Mills, Stefanie Zollmann, Jonathan VenturaIEEE VR 2020 · 2 citations
- Global Structure-from-Motion Meets Feedforward ReconstructionLinfei Pan, Johannes L. Schönberger, Marc PollefeysCVPR 2026 · 5 citations
- PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesDiantao Tu, Hainan Cui, Xianwei Zheng, Shuhan ShenCVPR 2024 · 4 citations
