EMR-MSF: Self-Supervised Recurrent Monocular Scene Flow Exploiting Ego-Motion Rigidity
Zijie Jiang, Masatoshi Okutomi
Abstract
Self-supervised monocular scene flow estimation, aiming to understand both 3D structures and 3D motions from two temporally consecutive monocular images, has received increasing attention for its simple and economical sensor setup. However, the accuracy of current methods suffers from the bottleneck of less-efficient network architecture and lack of motion rigidity for regularization. In this paper, we propose a superior model named EMR-MSF by borrowing the advantages of network architecture design under the scope of supervised learning. We further impose explicit and robust geometric constraints with an elaborately constructed ego-motion aggregation module where a rigidity soft mask is proposed to filter out dynamic regions for stable ego-motion estimation using static regions. Moreover, we propose a motion consistency loss along with a mask regularization loss to fully exploit static regions. Several efficient training strategies are integrated including a gradient detachment technique and an enhanced view synthesis process for better performance. Our proposed method outperforms the previous self-supervised works by a large margin and catches up to the performance of supervised methods. On the KITTI scene flow benchmark, our approach improves the SF-all metric of the state-of-the-art self-supervised monocular method by 44% and demonstrates superior performance across sub-tasks including depth and visual odometry, amongst other self-supervised single-task or multi-task methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da23e05e-8605-4369-a5f7-5fd661eaf569Cited by top-tier papers2
- OAMaskFlow: Occlusion-Aware Motion Mask for Scene FlowXiongfeng Peng, Zhihua Liu, Weiming Li, Yamin Mao et al.AAAI 2025
- Zero-Shot Monocular Scene Flow Estimation in the WildYiqing Liang, Abhishek Badki, Hang Su, James Tompkin et al.CVPR 2025
Builds on13
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Forget About the LiDAR: Self-Supervised Depth Estimators with MED Probability VolumesJuan Luis Gonzalez Bello, Munchurl KimNeurIPS 2020 · 99 citations
- SENSE: A Shared Encoder Network for Scene-Flow EstimationHuaizu Jiang, Deqing Sun, Varun Jampani, Zhaoyang Lv et al.ICCV 2019 · 86 citations
- Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic ScenesFabian Brickwedde, Steffen Abraham, Rudolf MesterICCV 2019 · 55 citations
- Exploiting Rigidity Constraints for LiDAR Scene Flow EstimationGuanting Dong, Yueyi Zhang, Hanlin Li, Xiaoyan Sun et al.CVPR 2022 · 30 citations
Related papers
- Self-Supervised Multi-Frame Monocular Scene FlowJunhwa Hur, Stefan RothCVPR 2021
- EffiScene: Efficient Per-Pixel Rigidity Inference for Unsupervised Joint Learning of Optical Flow, Depth, Camera Pose and Motion SegmentationYang Jiao, Trac D. Tran, Guangming ShiCVPR 2021
- SLIM: Self-Supervised LiDAR Scene Flow and Motion SegmentationStefan Andreas Baur, David Josef Emmerichs, Frank Moosmann, Peter Pinggera et al.ICCV 2021 · 110 citations
- Imposing Consistency for Optical Flow EstimationJisoo Jeong, Jamie Menjay Lin, Fatih Porikli, Nojun KwakCVPR 2022 · 40 citations
- Attentive and Contrastive Learning for Joint Depth and Motion Field EstimationSeokju Lee, François Rameau, Fei Pan, In So KweonICCV 2021 · 38 citations
