Depth-Guided Sparse Structure-from-Motion for Movies and TV Shows
Sheng Liu, Xiaohan Nie, Raffay Hamid
Abstract
Existing approaches for Structure from Motion (SfM) produce impressive 3-D reconstruction results especially when using imagery captured with large parallax. However, to create engaging video-content in movies and TV shows, the amount by which a camera can be moved while filming a particular shot is often limited. The resulting small-motion parallax between video frames makes standard geometry-based SfM approaches not as effective for movies and TV shows. To address this challenge, we propose a simple yet effective approach that uses single-frame depth-prior obtained from a pretrained network to significantly improve geometry-based SfM for our small-parallax setting. To this end, we first use the depth-estimates of the detected keypoints to reconstruct the point cloud and camera-pose for initial two-view reconstruction. We then perform depth-regularized optimization to register new images and triangulate the new points during incremental reconstruction. To comprehensively evaluate our approach, we introduce a new dataset (StudioSfM) consisting of 130 shots with 21K frames from 15 studio-produced videos that are manually annotated by a professional CG studio. We demonstrate that our approach: (a) significantly improves the quality of 3-D reconstruction for our small-parallax setting, (b) does not cause any degradation for data with large-parallax, and (c) maintains the generalizability and scalability of geometry-based sparse SfM. Our dataset can be obtained at https://github.com/amazon-researchlsmall-baseline-camera-tracking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1656aefd-365e-47b7-af80-3ba820768f13Cited by top-tier papers9
- EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual QueriesJinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao et al.ICCV 2023 · 26 citations
- AeroTraj: Trajectory Planning for Fast, and Accurate 3D Reconstruction Using a Drone-based LiDARFawad Ahmad, Christina Suyong Shin, Rajrup Ghosh, John D'Ambrosio et al.UbiComp 2023 · 7 citations
- Ego3DT: Tracking Every 3D Object in Ego-centric VideosShengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun et al.ACM MM 2024 · 6 citations
- Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge DistillationWeining Ren, Hongjun Wang, Xiao Tan, Kai HanNeurIPS 2025 · 5 citations
- ProvNeRF: Modeling per Point Provenance in NeRFs as a Stochastic FieldKiyohiro Nakayama, Mikaela Angelina Uy, Yang You, Ke Li et al.NeurIPS 2024 · 3 citations
Builds on7
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
- AdaBins: Depth Estimation Using Adaptive BinsShariq Farooq Bhat, Ibraheem Alhashim, Peter WonkaCVPR 2021
Related papers
- MP-SfM: Monocular Surface Priors for Robust Structure-from-MotionZador Pataki, Paul-Edouard Sarlin, Johannes L. Schönberger, Marc PollefeysCVPR 2025
- MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic VideosZhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang et al.CVPR 2025
- Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic ScenesFabian Brickwedde, Steffen Abraham, Rudolf MesterICCV 2019 · 55 citations
- PR-RRN: Pairwise-Regularized Residual-Recursive Networks for Non-rigid Structure-from-MotionHaitian Zeng, Yuchao Dai, Xin Yu, Xiaohan Wang et al.ICCV 2021 · 12 citations
- RESfM: Robust Deep Equivariant Structure from MotionFadi Khatib, Yoni Kasten, Dror Moran, Meirav Galun et al.ICLR 2025
