UrbanMoE: A Sparse Multi-Modal Mixture-of-Experts Framework for Multi-Task Urban Region Profiling
Pingping Liu, Jiamiao Liu, Zijian Zhang, Hao Miao, Qi Jiang, Qingliang Li, Qiuzhan Zhou, Irwin King
Abstract
Urban region profiling, the task of characterizing geographical areas, is crucial for urban planning and resource allocation. However, existing research in this domain faces two significant limitations. First, most methods are confined to single-task prediction, failing to capture the interconnected, multi-faceted nature of urban environments where numerous indicators are deeply correlated. Second, the field lacks a standardized experimental benchmark, which severely impedes fair comparison and reproducible progress. To address these challenges, we first establish a comprehensive benchmark for multi-task urban region profiling, featuring multi-modal features and a diverse set of strong baselines to ensure a fair and rigorous evaluation environment. Concurrently, we propose UrbanMoE, the first sparse multi-modal, multi-expert framework specifically architected to solve the multi-task challenge. Leveraging a sparse Mixture-of-Experts architecture, it dynamically routes multi-modal features to specialized sub-networks, enabling the simultaneous prediction of diverse urban indicators. We conduct extensive experiments on three real-world datasets within our benchmark, where UrbanMoE consistently demonstrates superior performance over all baselines. Further in-depth analysis validates the efficacy and efficiency of our approach, setting a new state-of-the-art and providing the community with a valuable tool for future research in urban analytics 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6cd127b-88eb-4be6-83d0-160d88eec9aaCited by top-tier papers1
Ask how each one uses itBuilds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen et al.WWW 2024 · 124 citations
- Beyond the First Law of Geography: Learning Representations of Satellite Imagery by Leveraging Point-of-InterestsYanxin Xi, Tong Li, Huandong Wang, Yong Li et al.WWW 2022 · 86 citations
- ReFound: Crafting a Foundation Model for Urban Region Understanding upon Language and Visual FoundationsCongxi Xiao, Jingbo Zhou, Yixiong Xiao, Jizhou Huang et al.KDD 2024 · 18 citations
- Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial TopologyMeng Chen, Zechen Li, Weiming Huang, Yongshun Gong et al.KDD 2024 · 13 citations
Related papers
- UrbanExpert: Task-Conditioned Multi-Modal Fusion via Semantic Expert Routing for Urban Socioeconomic PredictionZechen Li, Hongwei Jia, Weiming Huang, Kai Zhao et al.KDD 2026
- Reconciling Geospatial Prediction and Retrieval via Sparse RepresentationsYi Li, Yuanlong Chen, Weiming Huang, Xiaoli Li et al.NeurIPS 2025
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban ScenariosBaichuan Zhou, Haote Yang, Dairong Chen, Junyan Ye et al.AAAI 2025 · 34 citations
- MoST: A Foundation Model for Multi-modality Spatio-temporal Traffic PredictionRonghui Xu, Jihao Chen, Jindong Tian, Chenjuan Guo et al.KDD 2026
- Urban Region Embedding via Multi-View Contrastive PredictionZechen Li, Weiming Huang, Kai Zhao, Min Yang et al.AAAI 2024 · 44 citations
