SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
Yu Sheng, Jiajun Deng, Xinran Zhang, Yu Zhang, Bei Hua, Yanyong Zhang, Jianmin Ji
摘要
A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods extend this paradigm by associating each primitive with a compressed semantic feature vector. However, these methods have two major limitations: (a) the naively compressed feature compromises expressiveness, affecting the model's ability to capture fine-grained semantics, and (b) the pixel-wise primitive prediction introduces redundancy in overlapping areas, causing unnecessary memory overhead. To this end, we introduce SpatialSplat, a feedforward framework that produces redundancy-aware Gaussians and capitalizes on a dual-field semantic representation. Particularly, with the insight that primitives within the same instance exhibit high semantic consistency, we decompose the semantic representation into a coarse feature field that encodes uncompressed semantics with minimal primitives, and a fine-grained yet low-dimensional feature field that captures detailed inter-instance relationships. Moreover, we propose a selective Gaussian mechanism, which retains only essential Gaussians in the scene, effectively eliminating redundant primitives. Our proposed Spatialsplat learns accurate semantic information and detailed instances prior with more compact 3D Gaussians, making semantic 3D reconstruction more applicable. We conduct extensive experiments to evaluate our method, demonstrating a remarkable 60% reduction in scene representation parameters while achieving superior performance over state-of-the-art methods. The code is available at https://github.com/shengyuuu/SpatialSplat.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning 3D Representations for Spatial Intelligence from Unposed Multi-View ImagesBo Zhou, Qiuxia Lai, Zeren Sun, Xiangbo Shu 等CVPR 2026 · 被引用 1 次
- Learning Compact 3D Representations from Feed-Forward Novel View SynthesisHonggyu An, Jaewoo Jung, Mungyeom Kim, Chaehyun Kim 等CVPR 2026
- Learning Surgical Robotic Manipulation with 3D Spatial PriorsYu Sheng, Lidian Wang, Xiaomeng Chu, Jiajun Deng 等CVPR 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
相关 Paper
- AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric PriorsXiaoxue Zhang, Xiaoxu Zheng, Yixuan Yin, Tiao Zhao 等CVPR 2026 · 被引用 6 次
- SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned PredictionZicheng Zhang, Xiangting Meng, Ke Wu, Wenchao DingCVPR 2026 · 被引用 7 次
- GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse ViewsTianyu Chen, Wei Xiang, Kang Han, Yu Lu 等CVPR 2026 · 被引用 5 次
- PixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D ReconstructionDavid Charatan, Sizhe Lester Li, Andrea Tagliasacchi, Vincent SitzmannCVPR 2024
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse ViewsRanran Huang, Krystian MikolajczykICCV 2025 · 被引用 12 次
