MVGamba: Unify 3D Content Generation as State Space Sequence Modeling
Xuanyu Yi, Zike Wu, Qiuhong Shen, Qingshan Xu, Pan Zhou, Joo-Hwee Lim, Shuicheng Yan, Xinchao Wang, Hanwang Zhang
摘要
Recent 3D large reconstruction models (LRMs) can generate high-quality 3D content in sub-seconds by integrating multi-view diffusion models with scalable multi-view reconstructors. Current works further leverage 3D Gaussian Splatting as 3D representation for improved visual quality and rendering efficiency. However, we observe that existing Gaussian reconstruction models often suffer from multi-view inconsistency and blurred textures. We attribute this to the compromise of multi-view information propagation in favor of adopting powerful yet computationally intensive architectures (e.g., Transformers). To address this issue, we introduce MVGamba, a general and lightweight Gaussian reconstruction model featuring a multi-view Gaussian reconstructor based on the RNN-like State Space Model (SSM). Our Gaussian reconstructor propagates causal context containing multi-view information for cross-view self-refinement while generating a long sequence of Gaussians for fine-detail modeling with linear complexity. With off-the-shelf multi-view diffusion models integrated, MVGamba unifies 3D generation tasks from a single image, sparse images, or text prompts. Extensive experiments demonstrate that MVGamba outperforms state-of-the-art baselines in all 3D content generation scenarios with approximately only of the model size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Long-LRM: Long-Sequence Large Reconstruction Model for Wide-Coverage Gaussian SplatsZiwen Chen, Hao Tan, Kai Zhang, Sai Bi 等ICCV 2025 · 被引用 14 次
- tttLRM: Test-Time Training for Long Context and Autoregressive 3D ReconstructionChen Wang, Hao Tan, Wang Yifan, Zhiqin Chen 等CVPR 2026 · 被引用 11 次
- StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video StreamsZike Wu, Qi Yan, Xuanyu Yi, Lele Wang 等ICLR 2026 · 被引用 9 次
- Nautilus: Locality-Aware Autoencoder for Scalable Mesh GenerationYuxuan Wang, Xuanyu Yi, Haohan Weng, Qingshan Xu 等ICCV 2025 · 被引用 3 次
- StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image StreamsYang Li, Jinglu Wang, Lei Chu, Xiao Li 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
相关 Paper
- Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D GenerationQitong Yang, Mingtao Feng, Zijie Wu, Weisheng Dong 等CVPR 2025
- DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat GenerationChenguo Lin, Panwang Pan, Bangbang Yang, Zeming Li 等ICLR 2025
- iLRM: An Iterative Large 3D Reconstruction ModelGyeongjin Kang, Seungtae Nam, Seungkwon Yang, Xiangyu Sun 等CVPR 2026 · 被引用 19 次
- Generative Gaussian Splatting: Generating 3D Scenes with Video Diffusion PriorsKatja Schwarz, Norman Müller, Peter KontschiederICCV 2025 · 被引用 3 次
- Turbo3D: Ultra-fast Text-to-3D GenerationHanzhe Hu, Tianwei Yin, Fujun Luan, Yiwei Hu 等CVPR 2025
