Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language Sequences
Dingkang Yang, Haopeng Kuang, Shuai Huang, Lihua Zhang
Abstract
Understanding human behaviors and intents from videos is a challenging task. Video flows usually involve time-series data from different modalities, such as natural language, facial gestures, and acoustic information. Due to the variable receiving frequency for sequences from each modality, the collected multimodal streams are usually unaligned. For multimodal fusion of asynchronous sequences, the existing methods focus on projecting multiple modalities into a common latent space and learning the hybrid representations, which neglects the diversity of each modality and the commonality across different modalities. Motivated by this observation, we propose a Multimodal Fusion approach for learning modality-Specific and modality-Agnostic representations (MFSA) to refine multimodal representations and leverage the complementarity across different modalities. Specifically, a predictive self-attention module is used to capture reliable contextual dependencies and enhance the unique features over the modality-specific spaces. Meanwhile, we propose a hierarchical cross-modal attention module to explore the correlations between cross-modal elements over the modality-agnostic space. In this case, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, the modality-specific and -agnostic multimodal representations are used together for downstream tasks. Comprehensive experiments on three multimodal datasets clearly demonstrate the superiority of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e1e0b2c-4613-4ac6-be7a-471f13a046c1Cited by top-tier papers24
- How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent PerceptionDingkang Yang, Kun Yang, Yuzheng Wang, Jing Liu et al.NeurIPS 2023 · 160 citations
- Spatio-Temporal Domain Awareness for Multi-Agent Collaborative PerceptionKun Yang, Dingkang Yang, Jingyu Zhang, Mingcheng Li et al.ICCV 2023 · 99 citations
- DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisPan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen et al.AAAI 2025 · 84 citations
- The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Kexue Fu, Manning Wang et al.NeurIPS 2023 · 75 citations
- AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionDingkang Yang, Shuai Huang, Zhi Xu, Zhenpeng Li et al.ICCV 2023 · 72 citations
Builds on14
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 419 citations
- Context-Aware Emotion Recognition NetworksJiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park et al.ICCV 2019 · 285 citations
- CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating DeepfakesHao Huang, Yongtao Wang, Zhaoyu Chen, Yuze Zhang et al.AAAI 2022 · 131 citations
Related papers
- Attention is not Enough: Mitigating the Distribution Discrepancy in Asynchronous Multimodal Sequence FusionTao Liang, Guosheng Lin, Lei Feng, Yan Zhang et al.ICCV 2021 · 84 citations
- Progressive Modality Reinforcement for Human Multimodal Emotion Recognition From Unaligned Multimodal SequencesFengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan et al.CVPR 2021
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu et al.CVPR 2026 · 7 citations
- Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation LearningMingcheng Li, Dingkang Yang, Yang Liu, Shunli Wang et al.NeurIPS 2024 · 48 citations
- Cross-modality Representation Interactive Learning for Multimodal Sentiment AnalysisJian Huang, Yanli Ji, Yang Yang, Heng Tao ShenACM MM 2023 · 17 citations
