Transfer Learning with Synthetic Corpora for Spatial Role Labeling and Reasoning
Roshanak Mirzaee, Parisa Kordjamshidi
Abstract
Recent research shows synthetic data as a source of supervision helps pretrained language models (PLM) transfer learning to new target tasks/domains. However, this idea is less explored for spatial language. We provide two new data resources on multiple spatial language processing tasks. The first dataset is synthesized for transfer learning on spatial question answering (SQA) and spatial role labeling (SpRL). Compared to previous SQA datasets, we include a larger variety of spatial relation types and spatial expressions. Our data generation process is easily extendable with new spatial expression lexicons. The second one is a real-world SQA dataset with human-generated questions built on an existing corpus with SPRL annotations. This dataset can be used to evaluate spatial language processing models in realistic situations. We show pretraining with automatically generated data significantly improves the SOTA results on several SQA and SPRL benchmarks, particularly when the training data in the target domain is small.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang et al.ICLR 2024 · 176 citations
- Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language ModelsWenshan Wu, Shaoguang Mao, Yadong Zhang, Yan Xia et al.NeurIPS 2024 · 100 citations
- Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame BenchmarkFangjun Li, David C. Hogg, Anthony G. CohnAAAI 2024 · 60 citations
- CityGPT: Empowering Urban Spatial Cognition of Large Language ModelsJie Feng, Tianhui Liu, Yuwei Du, Siqi Guo et al.KDD 2025 · 9 citations
- USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban AgentsSiqi Lai, Yansong Ning, Zirui Yuan, Zhixi Chen et al.ICLR 2026 · 7 citations
Builds on4
- Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question AnsweringKaixin Ma, Filip Ilievski, Jonathan Francis, Yonatan Bisk et al.AAAI 2021 · 100 citations
- StepGame: A New Benchmark for Robust Multi-Hop Spatial Reasoning in TextsZhengxiang Shi, Qiang Zhang, Aldo LipaniAAAI 2022 · 100 citations
- Semi-Supervised Variational Reasoning for Medical Dialogue GenerationDongdong Li, Zhaochun Ren, Pengjie Ren, Zhumin Chen et al.SIGIR 2021 · 45 citations
- RuleBERT: Teaching Soft Rules to Pre-Trained Language ModelsMohammed Saeed, Naser Ahmadi, Preslav Nakov, Paolo PapottiEMNLP 2021 · 9 citations
Related papers
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic DataMichael Ogezi, Freda ShiACL 2025 · 19 citations
- Things not Written in Text: Exploring Spatial Commonsense from Visual SignalsXiao Liu, Da Yin, Yansong Feng, Dongyan ZhaoACL 2022
- Data-Centric Lessons To Improve Speech-Language PretrainingVishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang et al.ICLR 2026 · 3 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- Can Multimodal Large Language Models Understand Spatial Relations?Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou et al.ACL 2025 · 16 citations
