Feature-Level Collaboration: Joint Unsupervised Learning of Optical Flow, Stereo Depth and Camera Motion
Cheng Chi, Qingjie Wang, Tianyu Hao, Peng Guo, Xin Yang
Abstract
Precise estimation of optical flow, stereo depth and camera motion are important for the real-world 3D scene understanding and visual perception. Since the three tasks are tightly coupled with the inherent 3D geometric constraints, current studies have demonstrated that the three tasks can be improved through jointly optimizing geometric loss functions of several individual networks. In this paper, we show that effective feature-level collaboration of the networks for the three respective tasks could achieve much greater performance improvement for all three tasks than only loss-level joint optimization. Specifically, we propose a single network to combine and improve the three tasks. The network extracts the features of two consecutive stereo images, and simultaneously estimates optical flow, stereo depth and camera motion. The whole network mainly contains four parts: (I) a feature-sharing encoder to extract features of input images, which can enhance features' representation ability; (II) a pooled decoder to estimate both optical flow and stereo depth; (III) a camera pose estimation module which fuses optical flow and stereo depth information; (IV) a cost volume complement module to improve the performance of optical flow in static and occluded regions. Our method achieves state-of-the-art performance among the joint unsupervised methods, including optical flow and stereo depth estimation on KITTI 2012 and 2015 benchmarks, and camera motion estimation on KITTI VO dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8762304-d798-4667-a43c-715eb8e18bd2Cited by top-tier papers6
- Active Stereo Without Pattern ProjectorLuca Bartolomei, Matteo Poggi, Fabio Tosi, Andrea Conti et al.ICCV 2023 · 12 citations
- Federated Online Adaptation for Deep StereoMatteo Poggi, Fabio TosiCVPR 2024 · 11 citations
- Towards Open-World Generation of Stereo Images and Unsupervised MatchingFeng Qiao, Zhexiao Xiong, Eric Xing, Nathan JacobsICCV 2025 · 1 citation
- NeRF-Supervised Deep StereoFabio Tosi, Alessio Tonioni, Daniele De Gregorio, Matteo PoggiCVPR 2023
- Unsupervised Cumulative Domain Adaptation for Foggy Scene Optical FlowHanyu Zhou, Yi Chang, Wending Yan, Luxin YanCVPR 2023
Builds on5
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 265 citations
- Self-Supervised Monocular Scene Flow EstimationJunhwa Hur, Stefan RothCVPR 2020
- Distilled Semantics for Comprehensive Scene Understanding from VideosFabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi et al.CVPR 2020
- Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo MatchingPengpeng Liu, Irwin King, Michael R. Lyu, Jia XuCVPR 2020
Related papers
- SENSE: A Shared Encoder Network for Scene-Flow EstimationHuaizu Jiang, Deqing Sun, Varun Jampani, Zhaoyang Lv et al.ICCV 2019 · 86 citations
- EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion SegmentationYang Jiao, Trac D. Tran, Guangming ShiCVPR 2021
- RM-Depth: Unsupervised Learning of Recurrent Monocular Depth in Dynamic ScenesTak-Wai HuiCVPR 2022 · 62 citations
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual OdometryNan Yang, Lukas von Stumberg, Rui Wang, Daniel CremersCVPR 2020
- Two-in-One Depth: Bridging the Gap Between Monocular and Binocular Self-supervised Depth EstimationZhengming Zhou, Qiulei DongICCV 2023 · 16 citations
