Things not Written in Text: Exploring Spatial Commonsense from Visual Signals
Xiao Liu, Da Yin, Yansong Feng, Dongyan Zhao
Abstract
Spatial commonsense, the knowledge about spatial position and relationship between objects (like the relative size of a lion and a girl, and the position of a boy relative to a bicycle when cycling), is an important part of commonsense knowledge. Although pretrained language models (PLMs) succeed in many NLP tasks, they are shown to be ineffective in spatial commonsense reasoning. Starting from the observation that images are more likely to exhibit spatial commonsense than texts, we explore whether models with visual signals learn more spatial commonsense than text-based PLMs. We propose a spatial commonsense benchmark that focuses on the relative scales of objects, and the positional relationship between people and objects under different actions.We probe PLMs and models with visual signals, including vision-language pretrained models and image synthesis models, on this benchmark, and find that image synthesis models are more capable of learning accurate and consistent spatial knowledge than other models. The spatial knowledge from image synthesis models also helps in natural language understanding tasks that require spatial commonsense.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7397d94-21a9-420f-9184-d4c7dcad27c8Cited by top-tier papers9
- WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AIHai Dang, Frederik Brudy, George W. Fitzmaurice, Fraser AndersonUIST 2023 · 43 citations
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li et al.EMNLP 2022 · 27 citations
- Z-LaVI: Zero-Shot Language Solver Fueled by Visual ImaginationYue Yang, Wenlin Yao, Hongming Zhang, Xiaoyang Wang et al.EMNLP 2022 · 7 citations
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 4 citations
- Visually-augmented pretrained language models for NLP tasks without imagesHangyu Guo, Kun Zhou, Wayne Xin Zhao, Qinyu Zhang et al.ACL 2023 · 2 citations
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 198 citations
Related papers
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- SpaCE-Eval: A Benchmark for Real-World Multi-Modal ReasoningXuyou Yang, Yucheng Zhao, Wenxuan Zhang, Immanuel KohICLR 2026
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song et al.ICLR 2023 · 5 citations
- SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Modelsjing wu, Jianhua Wu, Jiayi Guan, Jiahong Chen et al.ICML 2026
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement LearningYuhong Liu, Beichen Zhang, Yuhang Zang, Yuhang Cao et al.CVPR 2026 · 43 citations
