Multi-View Stereo by Temporal Nonparametric Fusion
Yuxin Hou, Juho Kannala, Arno Solin
Abstract
We propose a novel idea for depth estimation from multiview image-pose pairs, where the model has capability to leverage information from previous latent-space encodings of the scene. This model uses pairs of images and poses, which are passed through an encoder-decoder model for disparity estimation. The novelty lies in soft-constraining the bottleneck layer by a nonparametric Gaussian process prior. We propose a pose-kernel structure that encourages similar poses to have resembling latent spaces. The flexibility of the Gaussian process (GP) prior provides adapting memory for fusing information from previous views. We train the encoder-decoder and the GP hyperparameters jointly end-to-end. In addition to a batch method, we derive a lightweight estimation scheme that circumvents standard pitfalls in scaling Gaussian process inference, and demonstrate how our scheme can run in real-time on smart devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ac783c4-798a-41b1-86c8-3da3e6c7d046Cited by top-tier papers29
- NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-view StereoYi Wei, Shaohui Liu, Yongming Rao, Wang Zhao et al.ICCV 2021 · 286 citations
- TransformerFusion: Monocular RGB Scene Reconstruction using TransformersAljaz Bozic, Pablo R. Palafox, Justus Thies, Angela Dai et al.NeurIPS 2021 · 185 citations
- VolumeFusion: Deep Depth Fusion for 3D Scene ReconstructionJaesung Choe, Sunghoon Im, François Rameau, Minjun Kang et al.ICCV 2021 · 83 citations
- Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object DetectionJinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer et al.ICLR 2023 · 71 citations
- PlanarRecon: Realtime 3D Plane Detection and Reconstruction from Posed Monocular VideosYiming Xie, Matheus Gadelha, Fengting Yang, Xiaowei Zhou et al.CVPR 2022 · 32 citations
Related papers
- A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose EmbeddingYitong Dong, Yijin Li, Zhaoyang Huang, Weikang Bian et al.NeurIPS 2024 · 7 citations
- PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric FusionZhiwei Zhang, Ruikai Xu, Weijian Zhang, Zhizhong Zhang et al.ACM MM 2025 · 2 citations
- Gaussian Process Priors for View-Aware InferenceYuxin Hou, Ari Heljakka, Arno SolinAAAI 2021 · 1 citation
- Mutual Adaptive Reasoning for Monocular 3D Multi-Person Pose EstimationJuze Zhang, Jingya Wang, Ye Shi, Fei Gao et al.ACM MM 2022 · 15 citations
- Augmenting Depth Estimation with Geospatial ContextScott Workman, Hunter BlantonICCV 2021 · 6 citations
