CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion Network
Jiajia Tang, Kang Li, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong
摘要
Multimodal sentiment analysis is the challenging research area that attends to the fusion of multiple heterogeneous modalities. The main challenge is the occurrence of some missing modalities during the multimodal fusion procedure. However, the existing techniques require all modalities as input, thus are sensitive to missing modalities at predicting time. In this work, the coupled-translation fusion network (CTFN) is firstly proposed to model bi-direction interplay via couple learning, ensuring the robustness in respect to missing modalities. Specifically, the cyclic consistency constraint is presented to improve the translation performance, allowing us directly to discard decoder and only embraces encoder of Transformer. This could contribute to a much lighter model. Due to the couple learning, CTFN is able to conduct bi-direction crossmodality intercorrelation parallelly. Based on CTFN, a hierarchical architecture is further established to exploit multiple bi-direction translations, leading to double multimodal fusing embeddings compared with traditional translation methods. Moreover, the convolution block is utilized to further highlight explicit interactions among those translations. For evaluation, CTFN was verified on two multimodal benchmarks with extensive ablation studies. The experiments demonstrate that the proposed framework achieves state-of-the-art or often competitive performance. Additionally, CTFN still maintains robustness when considering missing modality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Dynamically Adjust Word Representations Using Unaligned Multimodal InformationJiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding 等ACM MM 2022 · 被引用 59 次
- MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality SequencesWei Han, Hui Chen, Min-Yen Kan, Soujanya PoriaEMNLP 2022 · 被引用 13 次
- A Multi-Focus-Driven Multi-Branch Network for Robust Multimodal Sentiment AnalysisChuanqi Tao, Jiaming Li, Tianzi Zang, Peng GaoAAAI 2025 · 被引用 10 次
- Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing ModalitiesRui Liu, Haolin Zuo, Zheng Lian, Hongyu Yuan 等ACM MM 2025 · 被引用 6 次
- Whisper-UT: A Unified Translation Framework for Speech and TextCihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz 等EMNLP 2025
它引用的顶会 Paper1
相关 Paper
- Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Tianyi Liu, Jiantao ZhouSIGIR 2022 · 被引用 84 次
- Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation LearningMingcheng Li, Dingkang Yang, Yang Liu, Shunli Wang 等NeurIPS 2024 · 被引用 48 次
- Transformer-based Feature Reconstruction Network for Robust Multimodal Sentiment AnalysisZiqi Yuan, Wei Li, Hua Xu, Wenmeng YuACM MM 2021 · 被引用 186 次
- Mitigating Inconsistencies in Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Jiantao Zhou, Tianyi LiuEMNLP 2022 · 被引用 29 次
- A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing ModalitiesMingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang 等AAAI 2024 · 被引用 71 次
