Flash-Mono: Feed-Forward Accelerated Gaussian Splatting Monocular SLAM
Zicheng Zhang, Ke Wu, Xiangting Meng, Keyu Liu, Jieru Zhao, Wenchao Ding
摘要
Monocular 3D Gaussian Splatting SLAM suffers from critical limitations in time efficiency, geometric accuracy, and multi-view consistency. These issues stem from the time-consuming optimization and the lack of inter-frame scale consistency from single-frame geometry priors. We contend that a feed-forward paradigm, leveraging multi-frame context to predict Gaussian attributes directly, is crucial for addressing these challenges. We present Flash-Mono, a system composed of three core modules: a feed-forward prediction frontend, a 2D Gaussian Splatting mapping backend, and an efficient hidden-state-based loop closure module. We trained a recurrent feed-forward frontend model that progressively aggregates multi-frame visual features into a hidden state via cross attention and jointly predicts camera poses and per-pixel Gaussian properties. By directly predicting Gaussian attributes, our method bypasses the burdensome per-frame optimization required in optimization-based GS-SLAM, achieving a speedup while ensuring high-quality rendering. The power of our recurrent architecture extends beyond efficient prediction. The hidden states act as compact submap descriptors, facilitating efficient loop closure and global optimization to mitigate the long-standing challenge of drift. For enhanced geometric fidelity, we replace conventional 3D Gaussian ellipsoids with 2D Gaussian surfels. Extensive experiments demonstrate that Flash-Mono achieves state-of-the-art performance in both tracking and mapping quality, highlighting its potential for embodied perception and real-time reconstruction applications. Project page: https://victkk.github.io/flash-mono.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- 2D Gaussian Splatting for Geometrically Accurate Radiance FieldsBinbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger 等SIGGRAPH 2024 · 被引用 660 次
- Gaussian Splatting SLAMHidenobu Matsuki, Riku Murai, Paul H. J. Kelly, Andrew J. DavisonCVPR 2024 · 被引用 328 次
相关 Paper
- DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular VideosChieh Hubert Lin, Zhaoyang Lv, Songyin Wu, Zhen Xu 等NeurIPS 2025 · 被引用 15 次
- MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation ModelsYifan Liu, Keyu Fan, Weihao Yu, Chenxin Li 等CVPR 2025
- Splatter Image: Ultra-Fast Single-View 3D ReconstructionStanislaw Szymanowicz, Christian Rupprecht, Andrea VedaldiCVPR 2024 · 被引用 132 次
- REACT3D: Real-time Edge Accelerator for Incremental Training in 3D Gaussian Splatting based SLAM SystemsHongyi Wang, Zhenhua Zhu, Tianchen Zhao, Yunfei Xiang 等MICRO 2025 · 被引用 3 次
- Segs-Slam: Structure-Enhanced 3D Gaussian Splatting Slam With Appearance EmbeddingTianci Wen, Zhiang Liu, Yongchun FangICCV 2025 · 被引用 4 次
