SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language
Zehan Wang, Sashuai Zhou, Shaoxuan He, Haifeng Huang, Lihe Yang, Ziang Zhang, Xize Cheng, Shengpeng Ji, Tao Jin, Hengshuang Zhao, Zhou Zhao
摘要
Contrastive Language-Image Pre-training (CLIP) learns robust visual models through language supervision, making it a crucial visual encoding technique for various applications. However, CLIP struggles with comprehending spatial concepts in images, potentially restricting the spatial intelligence of CLIP-based AI systems. In this work, we propose SpatialCLIP, an enhanced version of CLIP with better spatial understanding capabilities. To capture the intricate 3D spatial relationships in images, we improve both "visual model" and "language supervision" of CLIP. Specifically, we design 3D-inspired ViT to replace the standard ViT in CLIP. By lifting 2D image tokens into 3D space and incorporating design insights from point cloud networks, our visual model gains greater potential for spatial perception. Meanwhile, captions with accurate and detailed spatial information are very rare. To explore better language supervision for spatial understanding, we re-caption images and perturb their spatial phrases as negative descriptions, which compels the visual model to seek spatial cues to distinguish these hard negative captions. With the enhanced visual model, we introduce SpatialLLaVA, following the same LLaVA-1.5 training protocol, to investigate the importance of visual representations for MLLM's spatial intelligence. Furthermore, we create SpatialBench, a benchmark specifically designed to evaluate CLIP and MLLM in spatial reasoning. Spatial-CLIP and SpatialLLaVA achieve substantial performance improvements, demonstrating stronger capabilities in spatial perception and reasoning, while maintaining comparable results on general-purpose benchmarks. Recently, Contrastive Language-Image Pretraining (CLIP) models [13, 22, 47, 53, 73] have demonstrated the ability to * Equal Contribution. † Corresponding author. Accurate: The darker-haired cat is behind the white cat. CLIP Score: 22.28 Category Error: The darker-haired bird is behind the white cat. CLIP Score: 20.43 (-1.85) Accurate: The white refrigerator is closer to the camera than brown door. CLIP Score: 29.74 Category Error: The white cabinet is closer to the camera brown door. CLIP Score: 25.66 (-4.08) Accurate: The dark coffee machine is to the left of the mug. CLIP Score: 22.10 Category Error: The ceramic cup is to the left of the mug. CLIP Score: 19.48 (-2.62) Accurate: Two people are standing behind a pile of huge-sized pipes. CLIP Score: 26.62 Category Error: Two people are standing behind a pile of huge-sized papers. CLIP Score: 16.40 (-10.22) Attribute Error: The light-haired cat is behind the white cat. CLIP Score: 21.22 (-1.06) Spatial Error: The darker-haired cat is to the right of the white cat. CLIP Score: 23.66 (+1.38) Attribute Error: The dark refrigerator is closer to the camera than opened door. CLIP Score: 29.39 (-0.35) Spatial Error: The white refrigerator is further from the camera than brown door. CLIP Score: 29.62 (-0.12) Attribute Error: The old broken coffee machine is to the left of the mug. CLIP Score: 21.27 (-0.83) Spatial Error: The dark coffee machine is in front of the mug. CLIP Score: 22.53 (+0.43) Attribute Error: Two people are standing behind a pile of small pipes. CLIP Score: 25.78 (-0.84) Spatial Error: Two people are standing inside a pile of huge-sized pipes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Textual Supervision Enhances Geospatial Representations in Vision-Language ModelsMarcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, Luiz Felipe Vecchietti 等ICML 2026
- SpaceCLIP: A Vision-Language Pretraining Framework With Spatial Reconstruction On TextBo Zou, Chao Yang, Chengbin Quan, Youjian ZhaoACM MM 2023 · 被引用 1 次
- Expand VSR Benchmark for VLLM to Expertize in Spatial RulesPeijin Xie, Lin Sun, Bingquan Liu, Dexin Wang 等AAAI 2025
- un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIPYinqi Li, Jiahe Zhao, Hong Chang, Ruibing Hou 等NeurIPS 2025 · 被引用 6 次
- Refining CLIP's Spatial Awareness: A Visual-Centric PerspectiveCongpei Qiu, Yanhao Wu, Wei Ke, Xiuxiu Bai 等ICLR 2025
