Deep Two-View Structure-From-Motion Revisited
Jianyuan Wang, Yiran Zhong, Yuchao Dai, Stan Birchfield, Kaihao Zhang, Nikolai Smolyanskiy, Hongdong Li
摘要
Two-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem by either recovering absolute pose scales from two consecutive frames or predicting a depth map from a single image, both of which are ill-posed problems. In contrast, we propose to revisit the problem of deep two-view SfM by leveraging the well-posedness of the classic pipeline. Our method consists of 1) an optical flow estimation network that predicts dense correspondences between two frames; 2) a normalized pose estimation module that computes relative camera poses from the 2D optical flow correspondences, and 3) a scale-invariant depth estimation network that leverages epipolar geometry to reduce the search space, refine the dense correspondences, and estimate relative depth maps. Extensive experiments show that our method outperforms all state-of-the-art two-view SfM methods by a clear margin on KITTI depth, KITTI VO, MVS, Scenes11, and SUN3D datasets in both relative pose and depth estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 被引用 158 次
- Implicit Motion Handling for Video Camouflaged Object DetectionXuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong 等CVPR 2022 · 被引用 83 次
- VGGSfM: Visual Geometry Grounded Deep Structure from MotionJianyuan Wang, Nikita Karaev, Christian Rupprecht, David NovotnýCVPR 2024 · 被引用 48 次
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao 等ACM MM 2022 · 被引用 23 次
- A Consistency-Aware Spot-Guided Transformer for Versatile and Hierarchical Point Cloud RegistrationRenlang Huang, Yufan Tang, Jiming Chen, Liang LiNeurIPS 2024 · 被引用 17 次
它引用的顶会 Paper5
- Hierarchical Neural Architecture Search for Deep Stereo MatchingXuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai 等NeurIPS 2020 · 被引用 436 次
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou 等ICCV 2019 · 被引用 373 次
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 被引用 314 次
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 被引用 265 次
- Displacement-Invariant Matching Cost Learning for Accurate Optical Flow EstimationJianyuan Wang, Yiran Zhong, Yuchao Dai, Kaihao Zhang 等NeurIPS 2020 · 被引用 83 次
相关 Paper
- Towards Better Generalization: Joint Depth-Pose Learning Without PoseNetWang Zhao, Shaohui Liu, Yezhi Shu, Yong-Jin LiuCVPR 2020
- PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesDiantao Tu, Hainan Cui, Xianwei Zheng, Shuhan ShenCVPR 2024 · 被引用 4 次
- LightedDepth: Video Depth Estimation in Light of Limited Inference View AnglesShengjie Zhu, Xiaoming LiuCVPR 2023
- Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?Lijun Wang, Yifan Wang, Linzhao Wang, Yunlong Zhan 等ICCV 2021 · 被引用 48 次
- Scale-flow: Estimating 3D Motion from VideoHan Ling, Quansen Sun, Zhenwen Ren, Yazhou Liu 等ACM MM 2022 · 被引用 7 次
