Attention is not Enough: Mitigating the Distribution Discrepancy in Asynchronous Multimodal Sequence Fusion
Tao Liang, Guosheng Lin, Lei Feng, Yan Zhang, Fengmao Lv
摘要
Videos flow as the mixture of language, acoustic, and vision modalities. A thorough video understanding needs to fuse time-series data of different modalities for prediction. Due to the variable receiving frequency for sequences from each modality, there usually exists inherent asynchrony across the collected multimodal streams. Towards an efficient multimodal fusion from asynchronous multimodal streams, we need to model the correlations between elements from different modalities. The recent Multimodal Transformer (MulT) approach extends the self-attention mechanism of the original Transformer network to learn the crossmodal dependencies between elements. However, the direct replication of self-attention will suffer from the distribution mismatch across different modality features. As a result, the learnt crossmodal dependencies can be unreliable. Motivated by this observation, this work proposes the Modality-Invariant Crossmodal Attention (MICA) approach towards learning crossmodal interactions over modality-invariant space in which the distribution mismatch between different modalities is well bridged. To this end, both the marginal distribution and the elements with high-confidence correlations are aligned over the common space of the query and key vectors which are computed from different modalities. Experiments on three standard benchmarks of multimodal video understanding clearly validate the superiority of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du 等ACM MM 2022 · 被引用 260 次
- Incomplete Multimodality-Diffused Emotion RecognitionYuanzhi Wang, Yong Li, Zhen CuiNeurIPS 2023 · 被引用 155 次
- DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisPan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen 等AAAI 2025 · 被引用 84 次
- Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language SequencesDingkang Yang, Haopeng Kuang, Shuai Huang, Lihua ZhangACM MM 2022 · 被引用 64 次
- Dynamically Adjust Word Representations Using Unaligned Multimodal InformationJiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding 等ACM MM 2022 · 被引用 59 次
它引用的顶会 Paper2
相关 Paper
- Multimodal Video Summarization via Time-Aware TransformersXindi Shang, Zehuan Yuan, Anran Wang, Changhu WangACM MM 2021 · 被引用 31 次
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan 等CVPR 2021
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang 等NeurIPS 2021 · 被引用 782 次
- Contextual Augmented Global Contrast for Multimodal Intent RecognitionKaili Sun, Zhiwen Xie, Mang Ye, Huyin ZhangCVPR 2024 · 被引用 19 次
- Everything at Once - Multi-modal Fusion Transformer for Video RetrievalNina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas 等CVPR 2022 · 被引用 4 次
