CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image
Wonseok Roh, Hwanhee Jung, Jong Wook Kim, Seunggwan Lee, Innfarn Yoo, Andreas Lugmayr, Seunggeun Chi, Karthik Ramani, Sangpil Kim
Abstract
Recently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by per-pixel 3D Gaussian primitives, from just a few images in a single forward pass. Unlike multi-view methods that benefit from cross-view correspondences, 3D scene reconstruction with a single-view image remains an underexplored area. In this work, we introduce CATSplat, a novel generalizable transformer-based framework designed to break through the inherent constraints in monocular settings. First, we propose leveraging textual guidance from a visual-language model to complement insufficient information from single-view image features. By incorporating scene-specific contextual details from text embeddings through cross-attention, we pave the way for context-aware 3D scene reconstruction beyond relying solely on visual cues. Moreover, we advocate utilizing spatial guidance from 3D point features toward comprehensive geometric understanding under monocular settings. With 3D priors, image features can capture rich structural insights for predicting 3D Gaussians without multi-view techniques. Extensive experiments on large-scale datasets demonstrate the state-of-the-art performance of CATSplat in single-view 3D scene reconstruction with high-quality novel view synthesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence ImagesGuichen Huang, Ruoyu Wang, Xiangjun Gao, Che Sun et al.AAAI 2026 · 6 citations
- PASTA: Part-Aware Sketch-to-3D Shape Generation with Text-Aligned PriorSeunggwan Lee, Hwanhee Jung, Byoungsoo Koh, Qixing Huang et al.ICCV 2025 · 2 citations
- Reconstruction Using the Invisible: Intuition from NIR and Metadata for Enhanced 3D Gaussian SplattingGyusam Chang, Tuan-Anh Vu, Vivek Alumootil, Harris Song et al.AAAI 2026 · 1 citation
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation ModelsYifan Liu, Keyu Fan, Weihao Yu, Chenxin Li et al.CVPR 2025
- TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian SplattingZhicong Wu, Hongbin Xu, Gang Xu, Ping Nie et al.ACM MM 2025 · 6 citations
- HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure PriorsPanwang Pan, Zhuo Su, Chenguo Lin, Zhen Fan et al.NeurIPS 2024 · 76 citations
- DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View ConsistencyTengfei Xiao, Yue Wu, Zhigang Gao, Yongzhe Yuan et al.AAAI 2026
- DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat GenerationChenguo Lin, Panwang Pan, Bangbang Yang, Zeming Li et al.ICLR 2025
