MM-Path: Multi-modal, Multi-granularity Path Representation Learning
Ronghui Xu, Hanyin Cheng, Chenjuan Guo, Hongfan Gao, Jilin Hu, Sean Bin Yang, Bin Yang
摘要
Developing effective path representations has become increasingly essential across various fields within intelligent transportation. Although pre-trained path representation learning models have shown improved performance, they predominantly focus on the topological structures from single modality data, i.e., road networks, overlooking the geometric and contextual features associated with path-related images, e.g., remote sensing images. Similar to human understanding, integrating information from multiple modalities can provide a more comprehensive view, enhancing both representation accuracy and generalization. However, variations in information granularity impede the semantic alignment of road network-based paths (road paths) and image-based paths (image paths), while the heterogeneity of multi-modal data poses substantial challenges for effective fusion and utilization. In this paper, we propose a novel Multi-modal, Multi-granularity Path Representation Learning Framework (MM-Path), which can learn a generic path representation by integrating modalities from both road paths and image paths. To enhance the alignment of multi-modal data, we develop a multi-granularity alignment strategy that systematically associates nodes, road sub-paths, and road paths with their corresponding image patches, ensuring the synchronization of both detailed local information and broader global contexts. To address the heterogeneity of multi-modal data effectively, we introduce a graph-based cross-modal residual fusion component designed to comprehensively fuse information across different modalities and granularities. Finally, we conduct extensive experiments on two large-scale real-world datasets under two downstream tasks, validating the effectiveness of the proposed MM-Path.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SSD-TS: Exploring the Potential of Linear State Space Models for Diffusion Models in Time Series ImputationHongfan Gao, Wangmeng Shen, Xiangfei Qiu, Ronghui Xu 等KDD 2025 · 被引用 5 次
- FlexiReg: Flexible Urban Region Representation LearningFengze Sun, Yanchuan Chang, Egemen Tanin, Shanika Karunasekera 等KDD 2025 · 被引用 3 次
- From Dialogue to Destination: Geography-Aware Large Language Models with Multimodal Fusion for Conversational RecommendationYeming Li, Chenxi Liu, Jie Zou, Cheng Long 等AAAI 2026 · 被引用 3 次
- Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?Shuo Liu, Di Yao, Yan Lin, Gao Cong 等KDD 2026 · 被引用 2 次
- TiWeaver: Unified Temporal Dynamics Modeling via Contextual PatchingZhe Li, Jindong Tian, Hao Miao, Zhi Lei 等KDD 2026 · 被引用 2 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsHangbo Bao, Wenhui Wang, Li Dong, Qiang Liu 等NeurIPS 2022 · 被引用 790 次
- Pathformer: Multi-scale Transformers with Adaptive Pathways for Time Series ForecastingPeng Chen, Yingying Zhang, Yunyao Cheng, Yang Shu 等ICLR 2024 · 被引用 197 次
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin 等ICCV 2023 · 被引用 152 次
相关 Paper
- Path-LLM: A Multi-Modal Path Representation Learning by Aligning and Fusing with Large Language ModelsYongfu Wei, Yan Lin, Hongfan Gao, Ronghui Xu 等WWW 2025 · 被引用 4 次
- Multimodal Trajectory Representation Learning for Travel Time EstimationZhi Liu, Xuyuan Hu, Xiao Han, Zhehao Dai 等WWW 2026
- A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal GraphsDongxiao He, AnKang Yang, Jitao Zhao, Di JinICML 2026
- Knowledge Graph Enhanced Multimodal Transformer for Image-Text RetrievalJuncheng Zheng, Meiyu Liang, Yang Yu, Yawen Li 等ICDE 2024 · 被引用 14 次
- Grid and Road Expressions Are Complementary for Trajectory Representation LearningSilin Zhou, Shuo Shang, Lisi Chen, Peng Han 等KDD 2025 · 被引用 7 次
