Visual Spatial Description: Controlled Spatial-Oriented Image-to-Text Generation
Yu Zhao, Jianguo Wei, Zhichao Lin, Yueheng Sun, Meishan Zhang, Min Zhang
Abstract
Image-to-text tasks, such as open-ended image captioning and controllable image description, have received extensive attention for decades. Here, we further advance this line of work by presenting Visual Spatial Description (VSD), a new perspective for image-to-text toward spatial semantics. Given an image and two objects inside it, VSD aims to produce one description focusing on the spatial perspective between the two objects. Accordingly, we manually annotate a dataset to facilitate the investigation of the newly-introduced task and build several benchmark encoder-decoder models by using VL-BART and VL-T5 as backbones. In addition, we investigate pipeline and joint end-to-end architectures for incorporating visual spatial relationship classification (VSRC) information into our model. Finally, we conduct experiments on our benchmark dataset to evaluate all our models. Results show that our models are impressive, providing accurate and human-like spatial-oriented text descriptions. Meanwhile, VSRC has great potential for VSD, and the joint end-to-end architecture is the better choice for their integration. We make the dataset and codes public for research purposes. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Generating Visual Spatial Description via Holistic 3D Scene UnderstandingYu Zhao, Hao Fei, Wei Ji, Jianguo Wei et al.ACL 2023 · 40 citations
- Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-imageYu Zhao, Hao Fei, Xiangtai Li, Libo Qin et al.NeurIPS 2024 · 2 citations
- Beyond Ranking: Fine-Grained Diagnostics and Self-Improvement for MLLMsMingze Xu, Zijing Zhao, Qiming Peng, Houwen Peng et al.ACL 2026
Builds on15
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
Related papers
- Enhancing Spatial Reasoning Through Visual and Textual ThinkingXun Liang, Xin Guo, Zhongming Jin, Weihang Pan et al.AAAI 2026
- Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched DescriptionXu Zhang, Jianzhong Huang, Lefei ZhangAAAI 2026 · 1 citation
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship DetectionShuo Yang, Yongqi Wang, Xiaofeng Ji, Xinxiao WuAAAI 2024 · 4 citations
- Can Transformers Capture Spatial Relations between Objects?Chuan Wen, Dinesh Jayaraman, Yang GaoICLR 2024 · 11 citations
