CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing
Tianhui Liu, Hetian Pang, Xin Zhang, Tianjian Ouyang, Zhiyuan Zhang, Jie Feng, Yong Li, Pan Hui
Abstract
Understanding urban socioeconomic conditions through visual data is a challenging yet essential task for sustainable urban development and policy planning. In this work, we introduce CityLens, a comprehensive benchmark designed to evaluate the capabilities of Large Vision-Language Models (LVLMs) in predicting socioeconomic indicators from satellite and street view imagery. We construct a multi-modal dataset covering a total of 17 globally distributed cities, spanning 6 key domains: economy, education, crime, transport, health, and environment, reflecting the multifaceted nature of urban life. Based on this dataset, we define 11 prediction tasks and utilize 3 evaluation paradigms: Direct Metric Prediction, Normalized Metric Estimation, and Feature-Based Regression. We benchmark 17 state-of-the-art LVLMs across these tasks. These make CityLens the most extensive socioeconomic benchmark to date in terms of geographic coverage, indicator diversity, and model scale. Our results reveal that while LVLMs demonstrate promising perceptual and reasoning capabilities, they still exhibit limitations in predicting urban socioeconomic indicators. CityLens provides a unified framework for diagnosing these limitations and guiding future efforts in using LVLMs to understand and predict urban socioeconomic patterns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- UrbanLLaVA: A Multi-Modal Large Language Model for Urban IntelligenceJie Feng, Shengyuan Wang, Tianhui Liu, Yanxin Xi et al.ICCV 2025 · 7 citations
- Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region ProfilingXixuan Hao, Yutian Jiang, Jiabo Liu, Yihang Yang et al.KDD 2026 · 2 citations
- UrbanMLLM: Joint Learning of Cross-view Imagery for Urban UnderstandingXin Zhang, Tianjian Ouyang, Yu Shang, Qingmin Liao et al.ICML 2026
Builds on14
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang et al.ICML 2024 · 306 citations
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen et al.WWW 2024 · 124 citations
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai et al.ICML 2024 · 110 citations
Related papers
- UrbanFeel:A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human PerspectiveJun He, Yi Lin, Zilong Huang, Jiacong Yin et al.ICLR 2026 · 7 citations
- Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial TopologyMeng Chen, Zechen Li, Weiming Huang, Yongshun Gong et al.KDD 2024 · 13 citations
- No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language ModelsAngéline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang et al.NeurIPS 2024 · 17 citations
- Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive EvaluationHong-Tao Yu, Yuxin Peng, Serge J. Belongie, Xiu-Shen WeiICLR 2026 · 21 citations
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long et al.ICLR 2026 · 8 citations
