DeepVideoMVS: Multi-View Stereo on Video With Recurrent Spatio-Temporal Fusion
Arda Düzçeker, Silvano Galliani, Christoph Vogel, Pablo Speciale, Mihai Dusmanu, Marc Pollefeys
Abstract
No Spatio-temporal Fusion (Backbone Network) Runtime 30 ms • FPS 34 • Memory = 795 MB (a) Groundtruth (c) With Spatio-temporal Fusion Runtime 37 ms • FPS 27 • Memory = 1061 MB Figure 1: 3D reconstructions of a scene from ScanNet [11]. Extending our stereo backbone with our proposed spatio-temporal fusion module improves the temporal consistency and accuracy of the predicted depth maps, leading to better reconstructions with negligible computational overhead. Runtime is per forward pass on an NVIDIA GTX 1080Ti with image size 320 × 256.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3219d92-c876-490a-8f33-70bffdf6b8bfCited by top-tier papers31
- TransformerFusion: Monocular RGB Scene Reconstruction using TransformersAljaz Bozic, Pablo R. Palafox, Justus Thies, Angela Dai et al.NeurIPS 2021 · 185 citations
- FreeSplat: Generalizable 3D Gaussian Splatting Towards Free View Synthesis of Indoor ScenesYunsong Wang, Tianxin Huang, Hanlin Chen, Gim Hee LeeNeurIPS 2024 · 112 citations
- VolumeFusion: Deep Depth Fusion for 3D Scene ReconstructionJaesung Choe, Sunghoon Im, François Rameau, Minjun Kang et al.ICCV 2021 · 83 citations
- Learning Attribute-driven Disentangled Representations for Interactive Fashion RetrievalYuxin Hou, Eleonora Vig, Michael Donoser, Loris BazzaniICCV 2021 · 58 citations
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang et al.NeurIPS 2022 · 50 citations
Builds on5
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- Exploiting Temporal Consistency for Real-Time Video Depth EstimationHaokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu et al.ICCV 2019 · 137 citations
- Multi-View Stereo by Temporal Nonparametric FusionYuxin Hou, Juho Kannala, Arno SolinICCV 2019 · 99 citations
- Softmax Splatting for Video Frame InterpolationSimon Niklaus, Feng LiuCVPR 2020
Related papers
- DG-Recon: Depth-Guided Neural 3D Scene ReconstructionJihong Ju, Ching Wei Tseng, Oleksandr Bailo, Georgi Dikov et al.ICCV 2023 · 21 citations
- NeuralRecon: Real-Time Coherent 3D Reconstruction From Monocular VideoJiaming Sun, Yiming Xie, Linghao Chen, Xiaowei Zhou et al.CVPR 2021
- A Decomposition Model for Stereo MatchingChengtang Yao, Yunde Jia, Huijun Di, Pengxiang Li et al.CVPR 2021
- DeepPruner: Learning Efficient Stereo Matching via Differentiable PatchMatchShivam Duggal, Shenlong Wang, Wei-Chiu Ma, Rui Hu et al.ICCV 2019 · 300 citations
- Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual OdometryZhaoxing Zhang, Junda Cheng, Gangwei Xu, Xiaoxiang Wang et al.AAAI 2025 · 9 citations
