UniLLM: A Unified Large Language Model for Multi?Modal Urban Dynamics Prediction
Yuhang Liu, Yingxue Zhang, Xin Zhang, Yanhua Li, Jun Luo
摘要
Modern cities generate vast streams of urban dynamics data reflecting mobility demand, environmental conditions, and traffic patterns. The value of these data lies not only in individual modalities but in their integration—urban signals are highly interdependent, with changes in one modality often influencing others. Consequently, predicting any single urban dynamic requires information from multiple interrelated sources. Although numerous methods—ranging from deep learning models to recent LLM-based approaches—have been proposed, most are limited in scope. They either focus on single-modality prediction, rely on rigid model designs that lack flexibility, or overlook inter-modal dependencies. As a result, they struggle to adapt to dynamic urban conditions and suffer from degraded predictive performance across modalities. In this paper, we propose UniLLM, a unified large language model for multi-modal urban dynamics prediction. At its core, UniLLM introduces a Unified Cross-Modal Alignment Module that transforms heterogeneous urban data into latent representations while preserving modality-specific patterns and capturing cross-modal correlations through a contrastive learning objective. To support dynamic adaptation across tasks and modalities, we design a Routing-Aware Prompting Mechanism that learns soft prompts based on task context and modality semantics. Furthermore, a Multi-Modal Memory-Guided Adaptive Algorithm employs replay-based gradient coordination and Frank–Wolfe optimization to mitigate cross-modal catastrophic forgetting during fine-tuning. Extensive experiments across multiple cities and urban modalities demonstrate that UniLLM consistently outperforms state-of-the-art baselines. These results highlight UniLLM's potential as a flexible and robust forecasting model for real-world, multi-modal urban environments.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language ModelsYuhang Liu, Yingxue Zhang, Xin Zhang, Ling Tian 等KDD 2025 · 被引用 3 次
- UrbanMLLM: Joint Learning of Cross-view Imagery for Urban UnderstandingXin Zhang, Tianjian Ouyang, Yu Shang, Qingmin Liao 等ICML 2026
- UrbanLLaVA: A Multi-Modal Large Language Model for Urban IntelligenceJie Feng, Shengyuan Wang, Tianhui Liu, Yanxin Xi 等ICCV 2025 · 被引用 7 次
- TransLLM: A Unified Multi-Task Large Language Model for Urban Transportation via Learnable PromptingJiaming Leng, Yunying Bi, Chuan Qin, Zhenya Huang 等ACL 2026
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin 等KDD 2024 · 被引用 75 次
