UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language Models
Yuhang Liu, Yingxue Zhang, Xin Zhang, Ling Tian, Yanhua Li, Jun Luo
Abstract
Understanding and predicting urban dynamics is crucial for managing transportation systems, optimizing urban planning, and enhancing public services. While neural network-based approaches have achieved success, they often rely on task-specific architectures and large volumes of data, limiting their ability to generalize across diverse urban scenarios. Meanwhile, Large Language Models (LLMs) offer strong reasoning and generalization capabilities, yet their application to spatial-temporal urban dynamics remains underexplored. Existing LLM-based methods struggle to effectively integrate multifaceted spatial-temporal data and fail to address distributional shifts between training and testing data, limiting their predictive reliability in real-world applications. To bridge this gap, we propose UrbanMind, a novel spatial-temporal LLM framework for multifaceted urban dynamics prediction that ensures both accurate forecasting and robust generalization. At its core, UrbanMind introduces Muffin-MAE, a multifaceted fusion masked autoencoder with specialized masking strategies that capture intricate spatial-temporal dependencies and intercorrelations among multifaceted urban dynamics. Additionally, we design a semantic-aware prompting and fine-tuning strategy that encodes spatial-temporal contextual details into prompts, enhancing LLMs' ability to reason over spatial-temporal patterns. To further improve generalization, we introduce a test time adaptation mechanism with a test data reconstructor, enabling UrbanMind to dynamically adjust to unseen test data by reconstructing LLM-generated embeddings. Extensive experiments on real-world urban dynamics datasets from multiple cities demonstrate the effectiveness of UrbanMind. The results consistently show that UrbanMind outperforms state-of-the-art baselines, achieving superior accuracy and strong generalization, even in zero-shot scenarios with no prior data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55e6fc62-9f11-479f-99f6-5ef2fc1995d7Cited by top-tier papers2
- AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban SensingXusen Guo, Mingxing Peng, Xixuan Hao, Xingchen Zou et al.WWW 2026 · 2 citations
- CLaDMoP: Learning Transferrable Models from Successful Clinical Trials via LLMsYiqing Zhang, Xiaozhong Liu, Fabricio MuraiKDD 2025 · 1 citation
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- GMAN: A Graph Multi-Attention Network for Traffic PredictionChuanpan Zheng, Xiaoliang Fan, Cheng Wang, Jianzhong QiAAAI 2020 · 1,858 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
Related papers
- UniLLM: A Unified Large Language Model for Multi?Modal Urban Dynamics PredictionYuhang Liu, Yingxue Zhang, Xin Zhang, Yanhua Li et al.KDD 2026
- TransLLM: A Unified Multi-Task Large Language Model for Urban Transportation via Learnable PromptingJiaming Leng, Yunying Bi, Chuan Qin, Zhenya Huang et al.ACL 2026
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin et al.KDD 2024 · 75 citations
- SeMob: Semantic Synthesis for Dynamic Urban Mobility PredictionRunfei Chen, Shuyang Jiang, Wei HuangEMNLP 2025
- ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language ModelTong Zhao, Junping Du, Zhe Xue, Meiyu Liang et al.AAAI 2026
