VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
Kaichen Zhou, Changhao Chen, Bing Wang, Muhamad Risqi Utama Saputra, Niki Trigoni, Andrew Markham
摘要
Recent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches towards deep fusion do not perform significantly better than models employing a single modality. We conjecture that this is because of the naive approaches to feature space fusion through summation or concatenation which do not take into account the different strengths of each modality. To address this, we propose an end-to-end framework, termed VM-Loc, to fuse different sensor inputs into a common latent space through a variational Product-of-Experts (PoE) followed by attention-based fusion. Unlike previous multimodal variational works directly adapting the objective function of vanilla variational auto-encoder, we show how camera localization can be accurately estimated through an unbiased objective function based on importance weighting. Our model is extensively evaluated on RGB-D datasets and the results prove the efficacy of our model. The source code is available at https://github.com/kaichen-z/VMLoc .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DynPoint: Dynamic Neural Point For View SynthesisKaichen Zhou, Jia-Xing Zhong, Sangyun Shin, Kai Lu 等NeurIPS 2023 · 被引用 46 次
- Text to Point Cloud Localization with Multi-Level Negative Contrastive LearningDunqiang Liu, Shujun Huang, Wen Li, Siqi Shen 等AAAI 2025 · 被引用 7 次
- MINIMA: Modality Invariant Image MatchingJiangwei Ren, Xingyu Jiang, Zizhuo Li, Dingkang Liang 等CVPR 2025
它引用的顶会 Paper3
- AtLoc: Attention Guided Camera LocalizationBing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao 等AAAI 2020 · 被引用 189 次
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi 等ICCV 2019 · 被引用 163 次
- Aligning Latent Spaces for 3D Hand Pose EstimationLinlin Yang, Shile Li, Dongheui Lee, Angela YaoICCV 2019 · 被引用 93 次
相关 Paper
- V-FUSE: Volumetric Depth Map Fusion with Long-Range ConstraintsNathaniel Burgdorfer, Philippos MordohaiICCV 2023 · 被引用 1 次
- Parameter-Efficient Variational AutoEncoder for Multimodal Multi-Interest RecommendationNhu-Thuat Tran, Hady W. LauwACM MM 2025
- Towards Visual Query Localization in the 3D WorldLiang Peng, Bohan Tan, Zhipeng Zhang, Haobo Li 等CVPR 2026 · 被引用 1 次
- Fusion Is Not Enough: Single Modal Attacks on Fusion Models for 3D Object DetectionZhiyuan Cheng, Hongjun Choi, Shiwei Feng, James Chenhao Liang 等ICLR 2024 · 被引用 32 次
- OmniVGGT: Omni-Modality Driven Visual Geometry Grounded TransformerHaosong Peng, Hao Li, Yalun Dai, Yushi Lan 等CVPR 2026 · 被引用 22 次
