GeoLM: Empowering Language Models for Geospatially Grounded Language Understanding
Zekun Li, Wenxuan Zhou, Yao-Yi Chiang, Muhao Chen
摘要
Humans subconsciously engage in geospatial reasoning when reading articles. We recognize place names and their spatial relations in text and mentally associate them with their physical locations on Earth. Although pretrained language models can mimic this cognitive process using linguistic context, they do not utilize valuable geospatial information in large, widely available geographical databases, e.g., OpenStreetMap. This paper introduces GEOLM ( ), a geospatially grounded language model that enhances the understanding of geo-entities in natural language. GEOLM leverages geo-entity mentions as anchors to connect linguistic information in text corpora with geospatial information extracted from geographical databases. GEOLM connects the two types of context through contrastive learning and masked language modeling. It also incorporates a spatial coordinate embedding mechanism to encode distance and direction relations to capture geospatial context. In the experiment, we demonstrate that GEOLM exhibits promising capabilities in supporting toponym recognition, toponym linking, relation extraction, and geo-entity typing, which bridge the gap between natural language processing and geospatial sciences. The code is publicly available at https://github.com/ knowledge-computing/geolm .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph ConstructionYansong Ning, Hao LiuNeurIPS 2024 · 被引用 35 次
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
- CityGPT: Empowering Urban Spatial Cognition of Large Language ModelsJie Feng, Tianhui Liu, Yuwei Du, Siqi Guo 等KDD 2025 · 被引用 9 次
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang 等ICLR 2025 · 被引用 6 次
- MoRA: Mobility as the Backbone for Geospatial Representation Learning at ScaleYa Wen, Jixuan Cai, Qiyao Ma, Linyan Li 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda 等EMNLP 2020 · 被引用 562 次
相关 Paper
- MGeo: Multi-Modal Geographic Language Model Pre-TrainingRuixue Ding, Boli Chen, Pengjun Xie, Fei Huang 等SIGIR 2023 · 被引用 29 次
- OsmT: Bridging Openstreetmap Queries and Natural Language With Open-Source Tag-Aware Language ModelsZhuoyue Wan, Wentao Hu, Chen Jason Zhang, Yuanfeng Song 等ICDE 2026
- GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsRohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke 等ICLR 2024 · 被引用 104 次
- Learning from Context or Names? An Empirical Study on Neural Relation ExtractionHao Peng, Tianyu Gao, Xu Han, Yankai Lin 等EMNLP 2020 · 被引用 185 次
- Geospatial Entity ResolutionPasquale Balsebre, Dezhong Yao, Gao Cong, Zhen HaiWWW 2022 · 被引用 21 次
