MTP: Exploring Multimodal Urban Traffic Profiling with Modality Augmentation and Spectrum Fusion
Haolong Xiang, Peisi Wang, Xiaolong Xu, Kun Yi, Xuyun Zhang, Quan Z. Sheng, Amin Beheshti, Wei Fan
Abstract
With rapid urbanization in the modern era, traffic signals from various sensors have been playing a significant role in monitoring the states of cities, which provides a strong foundation in ensuring safe travel, reducing traffic congestion and optimizing urban mobility. Most existing methods for traffic time series modeling often rely on the original data modality, i.e., numerical direct readings from the sensors in cities. However, this unimodal approach overlooks the semantic information existing in multimodal heterogeneous urban data in different perspectives, which hinders a comprehensive understanding of traffic signals and limits the accurate prediction of complex traffic dynamics. To address this problem, we propose a novel Multimodal framework, MTP, for urban Traffic Profiling, which learns multimodal features through numeric, visual, and textual perspectives in the frequency domain. The three branches drive a multimodal perspective of traffic signal learning for augmentation, while the frequency learning strategies delicately refine the information for extraction. Specifically, we first conduct the visual augmentation for the traffic time series, which transforms the original modality into periodicity images and frequency images for visual learning. Also, we augment descriptive texts for the traffic time series based on the specific topic, background information and item description for textual learning. To complement the numeric information, we utilize frequency multilayer perceptrons for learning on the original modality. We design a hierarchical contrastive learning on the three branches to fuse the three modalities. Finally, extensive experiments on six real-world datasets demonstrate superior performance compared with the state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72daeb2a-3149-45b1-98a2-981a12cc90a1Builds on23
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 898 citations
- Frequency-domain MLPs are More Effective Learners in Time Series ForecastingKun Yi, Qi Zhang, Wei Fan, Shoujin Wang et al.NeurIPS 2023 · 567 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationXinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li et al.WWW 2022 · 490 citations
- Hierarchical Multi-modal Contextual Attention Network for Fake News DetectionShengsheng Qian, Jinguang Wang, Jun Hu, Quan Fang et al.SIGIR 2021 · 273 citations
Related papers
- Urban Region Embedding via Multi-View Contrastive PredictionZechen Li, Weiming Huang, Kai Zhao, Min Yang et al.AAAI 2024 · 44 citations
- Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed InteractionShiyan Hu, Jianxin Jin, Yang Shu, Peng Chen et al.ICLR 2026 · 7 citations
- Role Hypergraph Contrastive Learning for Multivariate Time-Series AnalysisRundong Xue, Hao Hu, Zhitao Zeng, Xiangmin Han et al.AAAI 2026 · 1 citation
- FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent SpaceShengzhong Liu, Tomoyoshi Kimura, Dongxin Liu, Ruijie Wang et al.NeurIPS 2023 · 72 citations
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 138 citations
