Attentive and Contrastive Learning for Joint Depth and Motion Field Estimation
Seokju Lee, François Rameau, Fei Pan, In So Kweon
摘要
Estimating the motion of the camera together with the 3D structure of the scene from a monocular vision system is a complex task that often relies on the so-called scene rigidity assumption. When observing a dynamic environment, this assumption is violated which leads to an ambiguity between the ego-motion of the camera and the motion of the objects. To solve this problem, we present a self-supervised learning framework for 3D object motion field estimation from monocular videos. Our contributions are two-fold. First, we propose a two-stage projection pipeline to explicitly disentangle the camera ego-motion and the object motions with dynamics attention module, called DAM. Specifically, we design an integrated motion model that estimates the motion of the camera and object in the first and second warping stages, respectively, controlled by the attention module through a shared motion encoder. Second, we propose an object motion field estimation through contrastive sample consensus, called CSAC, taking advantage of weak semantic prior (bounding box from an object detector) and geometric constraints (each object respects the rigid body motion model). Experiments on KITTI, Cityscapes, and Waymo Open Dataset demonstrate the relevance of our approach and show that our method outperforms state-of-the-art algorithms for the tasks of self-supervised monocular depth estimation, object motion segmentation, monocular scene flow estimation, and visual odometry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth EstimationYouhong Wang, Yunji Liang, Hao Xu, Shaohui Jiao 等AAAI 2024 · 被引用 60 次
- Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical ScenesYihong Sun, Bharath HariharanNeurIPS 2023 · 被引用 58 次
- From-Ground-To-Objects: Coarse-to-Fine Self-supervised Monocular Depth Estimation of Dynamic Objects with Ground Contact PriorJaeho Moon, Juan Luis Gonzalez Bello, Byeongjun Kwon, Munchurl KimCVPR 2024 · 被引用 12 次
- PPEA-Depth: Progressive Parameter-Efficient Adaptation for Self-Supervised Monocular Depth EstimationYue-Jiang Dong, Yuan-Chen Guo, Ying-Tian Liu, Fang-Lue Zhang 等AAAI 2024 · 被引用 9 次
- Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth EstimationHoang Chuong Nguyen, Tianyu Wang, José M. Álvarez, Miaomiao LiuCVPR 2024 · 被引用 5 次
它引用的顶会 Paper7
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 被引用 397 次
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 被引用 265 次
- Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection ConsistencySeokju Lee, Sunghoon Im, Stephen Lin, In So KweonAAAI 2021 · 被引用 107 次
- Self-Supervised Monocular Scene Flow EstimationJunhwa Hur, Stefan RothCVPR 2020
相关 Paper
- EMR-MSF: Self-Supervised Recurrent Monocular Scene Flow Exploiting Ego-Motion RigidityZijie Jiang, Masatoshi OkutomiICCV 2023 · 被引用 5 次
- Multi-Frame Self-Supervised Depth Estimation with Multi-Scale Feature Fusion in Dynamic ScenesJiquan Zhong, Xiaolin Huang, Xiao YuACM MM 2023 · 被引用 6 次
- AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic ScenesXuanang Gao, Xiongbin Wu, Zhiwei Ning, Runze Yang 等AAAI 2026
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 被引用 44 次
- Object Concepts Emerge from MotionHaoqian Liang, Xiaohui Wang, Zhichao Li, Ya Yang 等NeurIPS 2025
