Depth-Guided Sparse Structure-from-Motion for Movies and TV Shows
Sheng Liu, Xiaohan Nie, Raffay Hamid
摘要
Existing approaches for Structure from Motion (SfM) produce impressive 3-D reconstruction results especially when using imagery captured with large parallax. However, to create engaging video-content in movies and TV shows, the amount by which a camera can be moved while filming a particular shot is often limited. The resulting small-motion parallax between video frames makes standard geometry-based SfM approaches not as effective for movies and TV shows. To address this challenge, we propose a simple yet effective approach that uses single-frame depth-prior obtained from a pretrained network to significantly improve geometry-based SfM for our small-parallax setting. To this end, we first use the depth-estimates of the detected keypoints to reconstruct the point cloud and camera-pose for initial two-view reconstruction. We then perform depth-regularized optimization to register new images and triangulate the new points during incremental reconstruction. To comprehensively evaluate our approach, we introduce a new dataset (StudioSfM) consisting of 130 shots with 21K frames from 15 studio-produced videos that are manually annotated by a professional CG studio. We demonstrate that our approach: (a) significantly improves the quality of 3-D reconstruction for our small-parallax setting, (b) does not cause any degradation for data with large-parallax, and (c) maintains the generalizability and scalability of geometry-based sparse SfM. Our dataset can be obtained at https://github.com/amazon-researchlsmall-baseline-camera-tracking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual QueriesJinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao 等ICCV 2023 · 被引用 26 次
- AeroTraj: Trajectory Planning for Fast, and Accurate 3D Reconstruction Using a Drone-based LiDARFawad Ahmad, Christina Suyong Shin, Rajrup Ghosh, John D'Ambrosio 等UbiComp 2023 · 被引用 7 次
- Ego3DT: Tracking Every 3D Object in Ego-centric VideosShengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun 等ACM MM 2024 · 被引用 6 次
- Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge DistillationWeining Ren, Hongjun Wang, Xiao Tan, Kai HanNeurIPS 2025 · 被引用 5 次
- ProvNeRF: Modeling per Point Provenance in NeRFs as a Stochastic FieldKiyohiro Nakayama, Mikaela Angelina Uy, Yang You, Ke Li 等NeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper7
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 被引用 314 次
- AdaBins: Depth Estimation Using Adaptive BinsShariq Farooq Bhat, Ibraheem Alhashim, Peter WonkaCVPR 2021
相关 Paper
- MP-SfM: Monocular Surface Priors for Robust Structure-from-MotionZador Pataki, Paul-Edouard Sarlin, Johannes L. Schönberger, Marc PollefeysCVPR 2025
- MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic VideosZhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang 等CVPR 2025
- Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic ScenesFabian Brickwedde, Steffen Abraham, Rudolf MesterICCV 2019 · 被引用 55 次
- PR-RRN: Pairwise-Regularized Residual-Recursive Networks for Non-rigid Structure-from-MotionHaitian Zeng, Yuchao Dai, Xin Yu, Xiaohan Wang 等ICCV 2021 · 被引用 12 次
- RESfM: Robust Deep Equivariant Structure from MotionFadi Khatib, Yoni Kasten, Dror Moran, Meirav Galun 等ICLR 2025
