MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time Series
Payal Mohapatra, Yueyuan Sui, Akash Pandey, Stephen Xia, Qi Zhu
摘要
From clinical healthcare to daily living, continuous sensor monitoring across multiple modalities has shown great promise for real-world intelligent decision-making but also faces various challenges. In this work, we introduce MAESTRO, a novel framework that overcomes key limitations of existing multimodal learning approaches: (1) reliance on a single primary modality for alignment, (2) pairwise modeling of modalities, and (3) assumption of complete modality observations. These limitations hinder the applicability of these approaches in real-world multimodal time-series settings, where primary modality priors are often unclear, the number of modalities can be large (making pairwise modeling impractical), and sensor failures often result in arbitrary missing observations. At its core, MAESTRO facilitates dynamic intra- and cross-modal interactions based on task relevance, and leverages symbolic tokenization and adaptive attention budgeting to construct long multimodal sequences, which are processed via sparse cross-modal attention. The resulting cross-modal tokens are routed through a sparse Mixture-of-Experts (MoE) mechanism, enabling black-box specialization under varying modality combinations. We evaluate MAESTRO against 10 baselines on four diverse datasets spanning three applications, and observe average relative improvements of 4% and 8% over the best existing multimodal and multivariate approaches, respectively, under complete observations. Under partial observations -- with up to 40% of missing modalities -- MAESTRO achieves an average 9% improvement. Further analysis also demonstrates the robustness and efficiency of MAESTRO's sparse, modality-aware design for learning from dynamic time series.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- iTransformer: Inverted Transformers Are Effective for Time Series ForecastingYong Liu, Tengge Hu, Haoran Zhang, Haixu Wu 等ICLR 2024 · 被引用 1,703 次
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
- TimesNet: Temporal 2D-Variation Modeling for General Time Series AnalysisHaixu Wu, Tengge Hu, Yong Liu, Hang Zhou 等ICLR 2023 · 被引用 423 次
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov 等AAAI 2021 · 被引用 393 次
相关 Paper
- Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-ExpertsSukwon Yun, Inyoung Choi, Jie Peng, Yangfan Wu 等NeurIPS 2024 · 被引用 98 次
- Taming Cascaded Mixture-of-Experts for Modality-missing Multi-modal Salient Object DetectionKunpeng Wang, Feifan Sun, Keke ChenAAAI 2026
- OmniField: Conditioned Neural Fields for Robust Multimodal Spatiotemporal LearningKevin Valencia, Thilina Balasooriya, Xihaier Luo, Shinjae Yoo 等ICLR 2026 · 被引用 1 次
- Building Massively Multimodal Foundation Models with Interaction-aware Mixture-of-ExpertsXing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang 等ICLR 2026 · 被引用 1 次
- Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential RecommendationShengzhe Zhang, Liyi Chen, Dazhong Shen, Chao Wang 等WWW 2025 · 被引用 29 次
