Diffusion Transformers as Open-World Spatiotemporal Foundation Models
Yuan Yuan, Chonghua Han, Jingtao Ding, Guozhen Zhang, Depeng Jin, Yong Li
摘要
The urban environment is characterized by complex spatio-temporal dynamics arising from diverse human activities and interactions. Effectively modeling these dynamics is essential for understanding and optimizing urban systems. In this work, we introduce UrbanDiT, a foundation model for open-world urban spatiotemporal learning that successfully scales up diffusion transformers in this field. UrbanDiT pioneers a unified model that integrates diverse data sources and types while learning universal spatio-temporal patterns across different cities and scenarios. This allows the model to unify both multi-data and multi-task learning, and effectively support a wide range of spatio-temporal applications. Its key innovation lies in the elaborated prompt learning framework, which adaptively generates both data-driven and task-specific prompts, guiding the model to deliver superior performance across various urban applications. UrbanDiT offers three advantages:
- It unifies diverse data types, such as grid-based and graph-based data, into a sequential format; 2) With task-specific prompts, it supports a wide range of tasks, including bi-directional spatio-temporal prediction, temporal interpolation, spatial extrapolation, and spatio-temporal imputation; and 3) It generalizes effectively to open-world scenarios, with its powerful zero-shot capabilities outperforming nearly all baselines with training data. UrbanDiT sets up a new benchmark for foundation models in the urban spatio-temporal domain. Code and datasets are publicly available at https://github.com/tsinghua-fib-lab/UrbanDiT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Adaptive Graph Convolutional Recurrent Network for Traffic ForecastingLei Bai, Lina Yao, Can Li, Xianzhi Wang 等NeurIPS 2020 · 被引用 2,206 次
相关 Paper
- UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal PredictionYuan Yuan, Jingtao Ding, Jie Feng, Depeng Jin 等KDD 2024 · 被引用 75 次
- UniSTD: Towards Unified Spatio-Temporal Learning across Diverse DisciplinesChen Tang, Xinzhu Ma, Encheng Su, Xiufeng Song 等CVPR 2025
- UrbanPG: An Efficient Framework with Personalized Context and General Backbone Interaction for Urban Spatio-Temporal LearningAoyu Liu, Yaying ZhangAAAI 2026
- UniLLM: A Unified Large Language Model for Multi?Modal Urban Dynamics PredictionYuhang Liu, Yingxue Zhang, Xin Zhang, Yanhua Li 等KDD 2026
- UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language ModelsYuhang Liu, Yingxue Zhang, Xin Zhang, Ling Tian 等KDD 2025 · 被引用 3 次
