MGeo: Multi-Modal Geographic Language Model Pre-Training
Ruixue Ding, Boli Chen, Pengjun Xie, Fei Huang, Xin Li, Qiang Zhang, Yao Xu
摘要
Query and point of interest (POI) matching is a core task in location-based services (LBS), e.g., navigation maps. It connects users' intent with real-world geographic information. Lately, pre-trained language models (PLMs) have made notable advancements in many natural language processing (NLP) tasks. To overcome the limitation that generic PLMs lack geographic knowledge for query-POI matching, related literature attempts to employ continued pre-training based on domain-specific corpus. However, a query generally describes the geographic context (GC) about its destination and contains mentions of multiple geographic objects like nearby roads and regions of interest (ROIs). These diverse geographic objects and their correlations are pivotal to retrieving the most relevant POI. Text-based single-modal PLMs can barely make use of the important GC and are therefore limited. In this work, we propose a novel method for query-POI matching, namely Multi-modal Geographic language model (MGeo), which comprises a geographic encoder and a multi-modal interaction module. Representing GC as a new modality, MGeo is able to fully extract multi-modal correlations to perform accurate query-POI matching. Moreover, there exists no publicly available query-POI matching benchmark. Intending to facilitate further research, we build a new open-source large-scale benchmark for this topic, i.e., Geographic TExtual Similarity (GeoTES). The POIs come from an open-source geographic information system (GIS) and the queries are manually generated by annotators to prevent privacy issues. Compared with several strong baselines, the extensive experiment results and detailed ablation analyses demonstrate that our proposed multi-modal geographic pre-training method can significantly improve the query-POI matching capability of PLMs with or without users' locations. Our code and benchmark are publicly available at https://github.com/PhantomGrapes/MGeo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen 等WWW 2024 · 被引用 124 次
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
- ReFound: Crafting a Foundation Model for Urban Region Understanding upon Language and Visual FoundationsCongxi Xiao, Jingbo Zhou, Yixiong Xiao, Jizhou Huang 等KDD 2024 · 被引用 18 次
- UrbanGeoEval: A City-Scale Benchmark for Evaluating Large Language Models in Geospatial ReasoningMutian Bao, Qiuyi Qi, Tian Liang, Jinjian Zhang 等ACL 2026
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 被引用 316 次
相关 Paper
- Geography-Aware Large Language Models for Next POI RecommendationWei Liu, Zhao Liu, Muzu Xie, Huaijie Zhu 等ICDE 2026 · 被引用 8 次
- GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingZekun Li, Wenxuan Zhou, Yao-Yi Chiang, Muhao ChenEMNLP 2023 · 被引用 25 次
- Mobility-Embedded POIs: Learning What A Place Is and How It Is Used from Human MovementMaria Despoina Siampou, Shushman Choudhury, Shang-Ling Hsu, Neha Arora 等ICML 2026 · 被引用 2 次
- POI-Enhancer: An LLM-based Semantic Enhancement Framework for POI Representation LearningJiawei Cheng, Jingyuan Wang, Yichuan Zhang, Jiahao Ji 等AAAI 2025 · 被引用 30 次
- GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal ModelsYushuo Zheng, Jiangyong Ying, Huiyu Duan, Chunyi Li 等AAAI 2026 · 被引用 2 次
