Towards Better Generalization: Joint Depth-Pose Learning Without PoseNet
Wang Zhao, Shaohui Liu, Yezhi Shu, Yong-Jin Liu
摘要
In this work, we tackle the essential problem of scale inconsistency for self-supervised joint depth-pose learning. Most existing methods assume that a consistent scale of depth and pose can be learned across all input samples, which makes the learning problem harder, resulting in degraded performance and limited generalization in indoor environments and long-sequence visual odometry application. To address this issue, we propose a novel system that explicitly disentangles scale from the network estimation. Instead of relying on PoseNet architecture, our method recovers relative pose by directly solving fundamental matrix from dense optical flow correspondence and makes use of a two-view triangulation module to recover an up-to-scale 3D structure. Then, we align the scale of the depth prediction with the triangulated point cloud and use the transformed depth map for depth error computation and dense reprojection check. Our whole system can be jointly trained end-to-end. Extensive experiments show that our system not only reaches state-of-the-art performance on KITTI depth and flow estimation, but also significantly improves the generalization ability of existing self-supervised depth-pose learning methods under a variety of challenging scenarios, and achieves state-of-the-art results among self-supervised learning-based methods on KITTI Odometry and NYUv2 dataset. Furthermore, we present some interesting findings on the limitation of PoseNet-based relative pose estimation methods in terms of generalization ability. Code is available at https://github.com/B1ueber2y/TrianFlow .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu 等ICCV 2023 · 被引用 380 次
- NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-view StereoYi Wei, Shaohui Liu, Yongming Rao, Wang Zhao 等ICCV 2021 · 被引用 286 次
- R-MSFM: Recurrent Multi-Scale Feature Modulation for Monocular Depth EstimatingZhongkai Zhou, Xinnan Fan, Pengfei Shi, Yuanxue XinICCV 2021 · 被引用 150 次
- Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth EstimationHyunyoung Jung, Eunhyeok Park, Sungjoo YooICCV 2021 · 被引用 133 次
- Regularizing Nighttime Weirdness: Efficient Self-supervised Monocular Depth Estimation in the DarkKun Wang, Zhenyu Zhang, Zhiqiang Yan, Xiang Li 等ICCV 2021 · 被引用 105 次
它引用的顶会 Paper5
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 被引用 397 次
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 被引用 265 次
- Learning Single Camera Depth Estimation Using Dual-PixelsRahul Garg, Neal Wadhwa, Sameer Ansari, Jonathan T. BarronICCV 2019 · 被引用 123 次
- Moving Indoor: Unsupervised Video Depth Learning in Challenging EnvironmentsJunsheng Zhou, Yuwang Wang, Kaihuai Qin, Wenjun ZengICCV 2019 · 被引用 74 次
相关 Paper
- Deep Two-View Structure-From-Motion RevisitedJianyuan Wang, Yiran Zhong, Yuchao Dai, Stan Birchfield 等CVPR 2021
- Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?Lijun Wang, Yifan Wang, Linzhao Wang, Yunlong Zhan 等ICCV 2021 · 被引用 48 次
- PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesDiantao Tu, Hainan Cui, Xianwei Zheng, Shuhan ShenCVPR 2024 · 被引用 4 次
- Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo MatchingPengpeng Liu, Irwin King, Michael R. Lyu, Jia XuCVPR 2020
- XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingLei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-BarICCV 2023 · 被引用 27 次
