Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial Topology
Meng Chen, Zechen Li, Weiming Huang, Yongshun Gong, Yilong Yin
Abstract
With the expansion and growth of cities, profiling urban areas with the advent of multi-modal urban datasets (e.g., points-of-interest and street view imagery) has become increasingly important in urban planing and management. Particularly, street view images have gained popularity for understanding the characteristics of urban areas due to its abundant visual information and inherent correlations with human activities. In this study, we define a street segment represented by multiple street view images as the minimum spatial unit for analysis and predict its functional and socioeconomic indicators, which presents several challenges in modeling spatial distributions of images on a street and the spatial topology (adjacency) of streets. Meanwhile, Large Language Models are capable of understanding imagery data based on its extraordinary knowledge base and unveil a remarkable opportunity for profiling streets with images. In view of the challenges and opportunity, we present a semi-supervised Urban Street Profiling Model (USPM) based on street view imagery and spatial adjacency of urban streets. Specifically, given a street with multiple images, we first employ a newly designed spatial context-based contrastive learning method to generate feature vectors of images and then apply the LSTM-based fusion method to encode multiple images on a street to yield the street visual representation; we then create the descriptions of street scenes for street view images based on the SPHINX (a large language model) and produce the street textual representation; finally, we build an urban street graph based on spatial topology (adjacency) and employ a semi-supervised graph learning algorithm to further encode the street representations for prediction. We conduct thorough experiments with real-world datasets to assess the proposed USPM. The experimental results demonstrate that USPM considerably outperforms baseline methods in two urban prediction tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c31d9f7a-37e9-494a-ba4e-b91623a7f5dcCited by top-tier papers8
- VecCity: A Taxonomy-guided Library for Map Entity Representation Learning [Experiment, Analysis & Benchmark]Wentao Zhang, Jingyuan Wang, Yifan Yang, Leong Hou UVLDB 2025 · 7 citations
- MM-Path: Multi-modal, Multi-granularity Path Representation LearningRonghui Xu, Hanyin Cheng, Chenjuan Guo, Hongfan Gao et al.KDD 2025 · 6 citations
- Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption SupervisionYimei Zhang, Guojiang Shen, Kaili Ning, Tongwei Ren et al.AAAI 2026 · 3 citations
- MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic PredictionMeng Chen, Junjie Yang, Zechen Li, Kai Zhao et al.ICML 2026
- SILO: Semantic Integration for Location Prediction with Large Language ModelsTianao Sun, Meng Chen, Bowen Zhang, Genan Dai et al.KDD 2025
Related papers
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen et al.WWW 2024 · 124 citations
- Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood EmbeddingZhecheng Wang, Haoyuan Li, Ram RajagopalAAAI 2020 · 113 citations
- UrbanMLLM: Joint Learning of Cross-view Imagery for Urban UnderstandingXin Zhang, Tianjian Ouyang, Yu Shang, Qingmin Liao et al.ICML 2026
- CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic SensingTianhui Liu, Hetian Pang, Xin Zhang, Tianjian Ouyang et al.ICLR 2026 · 10 citations
- Urban Region Embedding via Multi-View Contrastive PredictionZechen Li, Weiming Huang, Kai Zhao, Min Yang et al.AAAI 2024 · 44 citations
