Learning to Fuse Monocular and Multi-view Cues for Multi-frame Depth Estimation in Dynamic Scenes
Rui Li, Dong Gong, Wei Yin, Hao Chen, Yu Zhu, Kaixuan Wang, Xiaozhi Chen, Jinqiu Sun, Yanning Zhang
摘要
Multi-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to corrupted estimations. Many multi-frame methods handle dynamic areas by identifying them with explicit masks and compensating the multi-view cues with monocular cues represented as local monocular depth or features. The improvements are limited due to the uncontrolled quality of the masks and the underutilized benefits of the fusion of the two types of cues. In this paper, we propose a novel method to learn to fuse the multi-view and monocular cues encoded as volumes without needing the heuristically crafted masks. As unveiled in our analyses, the multiview cues capture more accurate geometric information in static areas, and the monocular cues capture more useful contexts in dynamic areas. To let the geometric perception learned from multi-view cues in static areas propagate to the monocular representation in dynamic areas and let monocular cues enhance the representation of multi-view cost volume, we propose a cross-cue fusion (CCF) module, which includes the cross-cue attention (CCA) to encode the spatially non-local relative intra-relations from each source to enhance the representation of the other. Experiments on real-world datasets prove the significant effectiveness and generalization ability of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- GoMVS: Geometrically Consistent Cost Aggregation for Multi-View StereoJiang Wu, Rui Li, Haofei Xu, Wenxun Zhao 等CVPR 2024 · 被引用 34 次
- From-Ground-To-Objects: Coarse-to-Fine Self-supervised Monocular Depth Estimation of Dynamic Objects with Ground Contact PriorJaeho Moon, Juan Luis Gonzalez Bello, Byeongjun Kwon, Munchurl KimCVPR 2024 · 被引用 12 次
- Improving the Convergence of Dynamic NeRFs via Optimal TransportSameera Ramasinghe, Violetta Shevchenko, Gil Avraham, Hisham Husain 等ICLR 2024 · 被引用 4 次
- Diving into the Fusion of Monocular Priors for Generalized Stereo MatchingChengtang Yao, Lidong Yu, Zhidan Liu, Jiaxi Zeng 等ICCV 2025 · 被引用 3 次
- MuGS: Multi-Baseline Generalizable Gaussian Splatting ReconstructionYaopeng Lou, Li Shen, Tianqi Liu, Jiaqi Li 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper18
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 被引用 487 次
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 被引用 397 次
- Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection ConsistencySeokju Lee, Sunghoon Im, Stephen Lin, In So KweonAAAI 2021 · 被引用 107 次
相关 Paper
- Crafting Monocular Cues and Velocity Guidance for Self-Supervised Multi-Frame Depth LearningXiaofeng Wang, Zheng Zhu, Guan Huang, Xu Chi 等AAAI 2023 · 被引用 31 次
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene UnderstandingXiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong LiCVPR 2026 · 被引用 2 次
- Multi-Frame Self-Supervised Depth Estimation with Multi-Scale Feature Fusion in Dynamic ScenesJiquan Zhong, Xiaolin Huang, Xiao YuACM MM 2023 · 被引用 6 次
- LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous DrivingQihao Sun, Jiarun Liu, Ziqian Ni, Jianyun Xu 等CVPR 2026 · 被引用 2 次
- SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth PriorsShaoqian Wang, Jiadai Sun, Bosen Hou, Qiang Wang 等CVPR 2026
