RoMo: Robust Motion Segmentation Improves Structure from Motion
Lily Goli, Sara Sabour, Mark J. Matthews, Marcus A. Brubaker, Dmitry Lagun, Alec Jacobson, David J. Fleet, Saurabh Saxena, Andrea Tagliasacchi
Abstract
There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38db137e-912a-4e77-908b-c43e37f727b4Cited by top-tier papers3
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng et al.ICCV 2025 · 29 citations
- Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger et al.ICCV 2025 · 11 citations
- RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGSChuanyu Fu, Yuqi Zhang, Kunbin Yao, Guanying Chen et al.ICCV 2025 · 6 citations
Builds on20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
Related papers
- ReFlow: Self-correction Motion Learning for Dynamic Scene ReconstructionYanzhe Liang, Ruijie Zhu, Hanzhi Chang, Zhuoyuan Li et al.CVPR 2026
- RGB-Only Supervised Camera Parameter Optimization in Dynamic ScenesFang Li, Hao Zhang, Narendra AhujaNeurIPS 2025
- MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion ScaffoldsJiahui Lei, Yijia Weng, Adam W. Harley, Leonidas J. Guibas et al.CVPR 2025
- Attentive and Contrastive Learning for Joint Depth and Motion Field EstimationSeokju Lee, François Rameau, Fei Pan, In So KweonICCV 2021 · 38 citations
- Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single VideoDavid Yifan Yao, Albert J. Zhai, Shenlong WangCVPR 2025
