SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language
Zehan Wang, Sashuai Zhou, Shaoxuan He, Haifeng Huang, Lihe Yang, Ziang Zhang, Xize Cheng, Shengpeng Ji, Tao Jin, Hengshuang Zhao, Zhou Zhao
Abstract
Contrastive Language-Image Pre-training (CLIP) learns robust visual models through language supervision, making it a crucial visual encoding technique for various applications. However, CLIP struggles with comprehending spatial concepts in images, potentially restricting the spatial intelligence of CLIP-based AI systems. In this work, we propose SpatialCLIP, an enhanced version of CLIP with better spatial understanding capabilities. To capture the intricate 3D spatial relationships in images, we improve both "visual model" and "language supervision" of CLIP. Specifically, we design 3D-inspired ViT to replace the standard ViT in CLIP. By lifting 2D image tokens into 3D space and incorporating design insights from point cloud networks, our visual model gains greater potential for spatial perception. Meanwhile, captions with accurate and detailed spatial information are very rare. To explore better language supervision for spatial understanding, we re-caption images and perturb their spatial phrases as negative descriptions, which compels the visual model to seek spatial cues to distinguish these hard negative captions. With the enhanced visual model, we introduce SpatialLLaVA, following the same LLaVA-1.5 training protocol, to investigate the importance of visual representations for MLLM's spatial intelligence. Furthermore, we create SpatialBench, a benchmark specifically designed to evaluate CLIP and MLLM in spatial reasoning. Spatial-CLIP and SpatialLLaVA achieve substantial performance improvements, demonstrating stronger capabilities in spatial perception and reasoning, while maintaining comparable results on general-purpose benchmarks. Recently, Contrastive Language-Image Pretraining (CLIP) models [13, 22, 47, 53, 73] have demonstrated the ability to * Equal Contribution. † Corresponding author. Accurate: The darker-haired cat is behind the white cat. CLIP Score: 22.28 Category Error: The darker-haired bird is behind the white cat. CLIP Score: 20.43 (-1.85) Accurate: The white refrigerator is closer to the camera than brown door. CLIP Score: 29.74 Category Error: The white cabinet is closer to the camera brown door. CLIP Score: 25.66 (-4.08) Accurate: The dark coffee machine is to the left of the mug. CLIP Score: 22.10 Category Error: The ceramic cup is to the left of the mug. CLIP Score: 19.48 (-2.62) Accurate: Two people are standing behind a pile of huge-sized pipes. CLIP Score: 26.62 Category Error: Two people are standing behind a pile of huge-sized papers. CLIP Score: 16.40 (-10.22) Attribute Error: The light-haired cat is behind the white cat. CLIP Score: 21.22 (-1.06) Spatial Error: The darker-haired cat is to the right of the white cat. CLIP Score: 23.66 (+1.38) Attribute Error: The dark refrigerator is closer to the camera than opened door. CLIP Score: 29.39 (-0.35) Spatial Error: The white refrigerator is further from the camera than brown door. CLIP Score: 29.62 (-0.12) Attribute Error: The old broken coffee machine is to the left of the mug. CLIP Score: 21.27 (-0.83) Spatial Error: The dark coffee machine is in front of the mug. CLIP Score: 22.53 (+0.43) Attribute Error: Two people are standing behind a pile of small pipes. CLIP Score: 25.78 (-0.84) Spatial Error: Two people are standing inside a pile of huge-sized pipes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87b5d5de-01ce-46d3-8ee6-a8fc0f31c8b7Cited by top-tier papers1
Ask how each one uses itBuilds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Textual Supervision Enhances Geospatial Representations in Vision-Language ModelsMarcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, Luiz Felipe Vecchietti et al.ICML 2026
- SpaceCLIP: A Vision-Language Pretraining Framework With Spatial Reconstruction On TextBo Zou, Chao Yang, Chengbin Quan, Youjian ZhaoACM MM 2023 · 1 citation
- Expand VSR Benchmark for VLLM to Expertize in Spatial RulesPeijin Xie, Lin Sun, Bingquan Liu, Dexin Wang et al.AAAI 2025
- un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIPYinqi Li, Jiahe Zhao, Hong Chang, Ruibing Hou et al.NeurIPS 2025 · 6 citations
- Refining CLIP's Spatial Awareness: A Visual-Centric PerspectiveCongpei Qiu, Yanhao Wu, Wei Ke, Xiuxiu Bai et al.ICLR 2025
