RoMo: Robust Motion Segmentation Improves Structure from Motion
Lily Goli, Sara Sabour, Mark J. Matthews, Marcus A. Brubaker, Dmitry Lagun, Alec Jacobson, David J. Fleet, Saurabh Saxena, Andrea Tagliasacchi
摘要
There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng 等ICCV 2025 · 被引用 29 次
- Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger 等ICCV 2025 · 被引用 11 次
- RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGSChuanyu Fu, Yuqi Zhang, Kunbin Yao, Guanying Chen 等ICCV 2025 · 被引用 6 次
它引用的顶会 Paper20
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
相关 Paper
- ReFlow: Self-correction Motion Learning for Dynamic Scene ReconstructionYanzhe Liang, Ruijie Zhu, Hanzhi Chang, Zhuoyuan Li 等CVPR 2026
- RGB-Only Supervised Camera Parameter Optimization in Dynamic ScenesFang Li, Hao Zhang, Narendra AhujaNeurIPS 2025
- MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion ScaffoldsJiahui Lei, Yijia Weng, Adam W. Harley, Leonidas J. Guibas 等CVPR 2025
- Attentive and Contrastive Learning for Joint Depth and Motion Field EstimationSeokju Lee, François Rameau, Fei Pan, In So KweonICCV 2021 · 被引用 38 次
- Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single VideoDavid Yifan Yao, Albert J. Zhai, Shenlong WangCVPR 2025
