Things not Written in Text: Exploring Spatial Commonsense from Visual Signals
Xiao Liu, Da Yin, Yansong Feng, Dongyan Zhao
摘要
Spatial commonsense, the knowledge about spatial position and relationship between objects (like the relative size of a lion and a girl, and the position of a boy relative to a bicycle when cycling), is an important part of commonsense knowledge. Although pretrained language models (PLMs) succeed in many NLP tasks, they are shown to be ineffective in spatial commonsense reasoning. Starting from the observation that images are more likely to exhibit spatial commonsense than texts, we explore whether models with visual signals learn more spatial commonsense than text-based PLMs. We propose a spatial commonsense benchmark that focuses on the relative scales of objects, and the positional relationship between people and objects under different actions.We probe PLMs and models with visual signals, including vision-language pretrained models and image synthesis models, on this benchmark, and find that image synthesis models are more capable of learning accurate and consistent spatial knowledge than other models. The spatial knowledge from image synthesis models also helps in natural language understanding tasks that require spatial commonsense.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AIHai Dang, Frederik Brudy, George W. Fitzmaurice, Fraser AndersonUIST 2023 · 被引用 43 次
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li 等EMNLP 2022 · 被引用 27 次
- Z-LaVI: Zero-Shot Language Solver Fueled by Visual ImaginationYue Yang, Wenlin Yao, Hongming Zhang, Xiaoyang Wang 等EMNLP 2022 · 被引用 7 次
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 被引用 4 次
- Visually-augmented pretrained language models for NLP tasks without imagesHangyu Guo, Kun Zhou, Wayne Xin Zhao, Qinyu Zhang 等ACL 2023 · 被引用 2 次
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi 等ICLR 2020 · 被引用 521 次
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 被引用 198 次
相关 Paper
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet 等NeurIPS 2024 · 被引用 166 次
- SpaCE-Eval: A Benchmark for Real-World Multi-Modal ReasoningXuyou Yang, Yucheng Zhao, Wenxuan Zhang, Immanuel KohICLR 2026
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song 等ICLR 2023 · 被引用 5 次
- SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Modelsjing wu, Jianhua Wu, Jiayi Guan, Jiahong Chen 等ICML 2026
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement LearningYuhong Liu, Beichen Zhang, Yuhang Zang, Yuhang Cao 等CVPR 2026 · 被引用 43 次
