MM-Path: Multi-modal, Multi-granularity Path Representation Learning
Ronghui Xu, Hanyin Cheng, Chenjuan Guo, Hongfan Gao, Jilin Hu, Sean Bin Yang, Bin Yang
Abstract
Developing effective path representations has become increasingly essential across various fields within intelligent transportation. Although pre-trained path representation learning models have shown improved performance, they predominantly focus on the topological structures from single modality data, i.e., road networks, overlooking the geometric and contextual features associated with path-related images, e.g., remote sensing images. Similar to human understanding, integrating information from multiple modalities can provide a more comprehensive view, enhancing both representation accuracy and generalization. However, variations in information granularity impede the semantic alignment of road network-based paths (road paths) and image-based paths (image paths), while the heterogeneity of multi-modal data poses substantial challenges for effective fusion and utilization. In this paper, we propose a novel Multi-modal, Multi-granularity Path Representation Learning Framework (MM-Path), which can learn a generic path representation by integrating modalities from both road paths and image paths. To enhance the alignment of multi-modal data, we develop a multi-granularity alignment strategy that systematically associates nodes, road sub-paths, and road paths with their corresponding image patches, ensuring the synchronization of both detailed local information and broader global contexts. To address the heterogeneity of multi-modal data effectively, we introduce a graph-based cross-modal residual fusion component designed to comprehensively fuse information across different modalities and granularities. Finally, we conduct extensive experiments on two large-scale real-world datasets under two downstream tasks, validating the effectiveness of the proposed MM-Path.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a451795a-4c20-41bd-8788-316c6853a04aCited by top-tier papers7
- SSD-TS: Exploring the Potential of Linear State Space Models for Diffusion Models in Time Series ImputationHongfan Gao, Wangmeng Shen, Xiangfei Qiu, Ronghui Xu et al.KDD 2025 · 5 citations
- FlexiReg: Flexible Urban Region Representation LearningFengze Sun, Yanchuan Chang, Egemen Tanin, Shanika Karunasekera et al.KDD 2025 · 3 citations
- From Dialogue to Destination: Geography-Aware Large Language Models with Multimodal Fusion for Conversational RecommendationYeming Li, Chenxi Liu, Jie Zou, Cheng Long et al.AAAI 2026 · 3 citations
- Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?Shuo Liu, Di Yao, Yan Lin, Gao Cong et al.KDD 2026 · 2 citations
- TiWeaver: Unified Temporal Dynamics Modeling via Contextual PatchingZhe Li, Jindong Tian, Hao Miao, Zhi Lei et al.KDD 2026 · 2 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsHangbo Bao, Wenhui Wang, Li Dong, Qiang Liu et al.NeurIPS 2022 · 790 citations
- Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series ForecastingPeng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu et al.ICLR 2024 · 197 citations
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin et al.ICCV 2023 · 152 citations
Related papers
- Path-LLM: A Multi-Modal Path Representation Learning by Aligning and Fusing with Large Language ModelsYongfu Wei, Yan Lin, Hongfan Gao, Ronghui Xu et al.WWW 2025 · 4 citations
- Multimodal Trajectory Representation Learning for Travel Time EstimationZhi Liu, Xuyuan Hu, Xiao Han, Zhehao Dai et al.WWW 2026
- A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal GraphsDongxiao He, AnKang Yang, Jitao Zhao, Di JinICML 2026
- Knowledge Graph Enhanced Multimodal Transformer for Image-Text RetrievalJuncheng Zheng, Meiyu Liang, Yang Yu, Yawen Li et al.ICDE 2024 · 14 citations
- Grid and Road Expressions Are Complementary for Trajectory Representation LearningSilin Zhou, Shuo Shang, Lisi Chen, Peng Han et al.KDD 2025 · 7 citations
