Unsupervised Collaborative Learning of Keyframe Detection and Visual Odometry Towards Monocular Deep SLAM
Lu Sheng, Dan Xu, Wanli Ouyang, Xiaogang Wang
Abstract
In this paper we tackle the joint learning problem of keyframe detection and visual odometry towards monocular visual SLAM systems. As an important task in visual SLAM, keyframe selection helps efficient camera relocalization and effective augmentation of visual odometry. To benefit from it, we first present a deep network design for the keyframe selection, which is able to reliably detect keyframes and localize new frames, then an end-to-end unsupervised deep framework further proposed for simultaneously learning the keyframe selection and the visual odometry tasks. As far as we know, it is the first work to jointly optimize these two complementary tasks in a single deep framework. To make the two tasks facilitate each other in the learning, a collaborative optimization loss based on both geometric and visual metrics is proposed. Extensive experiments on publicly available datasets (i.e. KITTI raw dataset and its odometry split [12] ) clearly demonstrate the effectiveness of the proposed approach, and new state-ofthe-art results are established on the unsupervised depth and pose estimation from monocular video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Leveraging Auxiliary Tasks with Affinity Learning for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd et al.ICCV 2021 · 152 citations
- SA-ConvONet: Sign-Agnostic Optimization of Convolutional Occupancy NetworksJiapeng Tang, Jiabao Lei, Dan Xu, Feiying Ma et al.ICCV 2021 · 84 citations
- Keyframe-Focused Visual Imitation LearningChuan Wen, Jierui Lin, Jianing Qian, Yang Gao et al.ICML 2021 · 27 citations
- Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-MotionGuoyu LuCVPR 2025
- Robust Consistent Video Depth EstimationJohannes Kopf, Xuejian Rong, Jia-Bin HuangCVPR 2021
Related papers
- Deep Two-View Structure-From-Motion RevisitedJianyuan Wang, Yiran Zhong, Yuchao Dai, Stan Birchfield et al.CVPR 2021
- Generalizing to the Open World: Deep Visual Odometry With Online AdaptationShunkai Li, Xin Wu, Yingdian Cao, Hongbin ZhaCVPR 2021
- Sequential Adversarial Learning for Self-Supervised Deep Visual OdometryShunkai Li, Fei Xue, Xin Wang, Zike Yan et al.ICCV 2019 · 58 citations
- Feature-Level Collaboration: Joint Unsupervised Learning of Optical Flow, Stereo Depth and Camera MotionCheng Chi, Qingjie Wang, Tianyu Hao, Peng Guo et al.CVPR 2021
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual OdometryNan Yang, Lukas von Stumberg, Rui Wang, Daniel CremersCVPR 2020
