UrbanMoE: A Sparse Multi-Modal Mixture-of-Experts Framework for Multi-Task Urban Region Profiling
Pingping Liu, Jiamiao Liu, Zijian Zhang, Hao Miao, Qi Jiang, Qingliang Li, Qiuzhan Zhou, Irwin King
摘要
Urban region profiling, the task of characterizing geographical areas, is crucial for urban planning and resource allocation. However, existing research in this domain faces two significant limitations. First, most methods are confined to single-task prediction, failing to capture the interconnected, multi-faceted nature of urban environments where numerous indicators are deeply correlated. Second, the field lacks a standardized experimental benchmark, which severely impedes fair comparison and reproducible progress. To address these challenges, we first establish a comprehensive benchmark for multi-task urban region profiling, featuring multi-modal features and a diverse set of strong baselines to ensure a fair and rigorous evaluation environment. Concurrently, we propose UrbanMoE, the first sparse multi-modal, multi-expert framework specifically architected to solve the multi-task challenge. Leveraging a sparse Mixture-of-Experts architecture, it dynamically routes multi-modal features to specialized sub-networks, enabling the simultaneous prediction of diverse urban indicators. We conduct extensive experiments on three real-world datasets within our benchmark, where UrbanMoE consistently demonstrates superior performance over all baselines. Further in-depth analysis validates the efficacy and efficiency of our approach, setting a new state-of-the-art and providing the community with a valuable tool for future research in urban analytics 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen 等WWW 2024 · 被引用 124 次
- Beyond the First Law of Geography: Learning Representations of Satellite Imagery by Leveraging Point-of-InterestsYanxin Xi, Tong Li, Huandong Wang, Yong Li 等WWW 2022 · 被引用 86 次
- ReFound: Crafting a Foundation Model for Urban Region Understanding upon Language and Visual FoundationsCongxi Xiao, Jingbo Zhou, Yixiong Xiao, Jizhou Huang 等KDD 2024 · 被引用 18 次
- Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial TopologyMeng Chen, Zechen Li, Weiming Huang, Yongshun Gong 等KDD 2024 · 被引用 13 次
相关 Paper
- UrbanExpert: Task-Conditioned Multi-Modal Fusion via Semantic Expert Routing for Urban Socioeconomic PredictionZechen Li, Hongwei Jia, Weiming Huang, Kai Zhao 等KDD 2026
- Reconciling Geospatial Prediction and Retrieval via Sparse RepresentationsYi Li, Yuanlong Chen, Weiming Huang, Xiaoli Li 等NeurIPS 2025
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosBaichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye 等AAAI 2025 · 被引用 34 次
- MoST: A Foundation Model for Multi-modality Spatio-temporal Traffic PredictionRonghui Xu, Jihao Chen, Jindong Tian, Chenjuan Guo 等KDD 2026
- Urban Region Embedding via Multi-View Contrastive PredictionZechen Li, Weiming Huang, Kai Zhao, Min Yang 等AAAI 2024 · 被引用 44 次
