Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood Embedding
Zhecheng Wang, Haoyuan Li, Ram Rajagopal
摘要
Understanding intrinsic patterns and predicting spatiotemporal characteristics of cities require a comprehensive representation of urban neighborhoods. Existing works relied on either inter- or intra-region connectivities to generate neighborhood representations but failed to fully utilize the informative yet heterogeneous data within neighborhoods. In this work, we propose Urban2Vec, an unsupervised multi-modal framework which incorporates both street view imagery and point-of-interest (POI) data to learn neighborhood embeddings. Specifically, we use a convolutional neural network to extract visual features from street view images while preserving geospatial similarity. Furthermore, we model each POI as a bag-of-words containing its category, rating, and review information. Analog to document embedding in natural language processing, we establish the semantic similarity between neighborhood (“document”) and the words from its surrounding POIs in the vector space. By jointly encoding visual, textual, and geospatial information into the neighborhood representation, Urban2Vec can achieve performances better than baseline models and comparable to fully-supervised methods in downstream prediction tasks. Extensive experiments on three U.S. metropolitan areas also demonstrate the model interpretability, generalization capability, and its value in neighborhood similarity analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen 等WWW 2024 · 被引用 124 次
- Beyond the First Law of Geography: Learning Representations of Satellite Imagery by Leveraging Point-of-InterestsYanxin Xi, Tong Li, Huandong Wang, Yong Li 等WWW 2022 · 被引用 86 次
- Urban Region Embedding via Multi-View Contrastive PredictionZechen Li, Weiming Huang, Kai Zhao, Min Yang 等AAAI 2024 · 被引用 44 次
- Urban Region Representation Learning with OpenStreetMap Building FootprintsYi Li, Weiming Huang, Gao Cong, Hao Wang 等KDD 2023 · 被引用 38 次
- BlockPlanner: City Block Generation with Vectorized Graph RepresentationLinning Xu, Yuanbo Xiangli, Anyi Rao, Nanxuan Zhao 等ICCV 2021 · 被引用 28 次
相关 Paper
- Multi-Scale Representation Learning for Spatial Feature Distributions using Grid CellsGengchen Mai, Krzysztof Janowicz, Bo Yan, Rui Zhu 等ICLR 2020 · 被引用 161 次
- Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial TopologyMeng Chen, Zechen Li, Weiming Huang, Yongshun Gong 等KDD 2024 · 被引用 13 次
- MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic PredictionMeng Chen, Junjie Yang, Zechen Li, Kai Zhao 等ICML 2026
- UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial RepresentationsDominik J. Mühlematter, Lin Che, Ye Hong, Martin Raubal 等ICML 2026
- ViCo: Word Embeddings From Visual Co-OccurrencesTanmay Gupta, Alexander G. Schwing, Derek HoiemICCV 2019 · 被引用 26 次
