CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion Network
Jiajia Tang, Kang Li, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong
Abstract
Multimodal sentiment analysis is the challenging research area that attends to the fusion of multiple heterogeneous modalities. The main challenge is the occurrence of some missing modalities during the multimodal fusion procedure. However, the existing techniques require all modalities as input, thus are sensitive to missing modalities at predicting time. In this work, the coupled-translation fusion network (CTFN) is firstly proposed to model bi-direction interplay via couple learning, ensuring the robustness in respect to missing modalities. Specifically, the cyclic consistency constraint is presented to improve the translation performance, allowing us directly to discard decoder and only embraces encoder of Transformer. This could contribute to a much lighter model. Due to the couple learning, CTFN is able to conduct bi-direction crossmodality intercorrelation parallelly. Based on CTFN, a hierarchical architecture is further established to exploit multiple bi-direction translations, leading to double multimodal fusing embeddings compared with traditional translation methods. Moreover, the convolution block is utilized to further highlight explicit interactions among those translations. For evaluation, CTFN was verified on two multimodal benchmarks with extensive ablation studies. The experiments demonstrate that the proposed framework achieves state-of-the-art or often competitive performance. Additionally, CTFN still maintains robustness when considering missing modality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a89dd64b-eced-4f64-8ad9-5f200b3e13faCited by top-tier papers7
- Dynamically Adjust Word Representations Using Unaligned Multimodal InformationJiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding et al.ACM MM 2022 · 59 citations
- MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality SequencesWei Han, Hui Chen, Min-Yen Kan, Soujanya PoriaEMNLP 2022 · 13 citations
- A Multi-Focus-Driven Multi-Branch Network for Robust Multimodal Sentiment AnalysisChuanqi Tao, Jiaming Li, Tianzi Zang, Peng GaoAAAI 2025 · 10 citations
- Hardness-Aware Dynamic Curriculum Learning for Robust Multimodal Emotion Recognition with Missing ModalitiesRui Liu, Haolin Zuo, Zheng Lian, Hongyu Yuan et al.ACM MM 2025 · 6 citations
- Whisper-UT: A Unified Translation Framework for Speech and TextCihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz et al.EMNLP 2025
Builds on1
Related papers
- Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Tianyi Liu, Jiantao ZhouSIGIR 2022 · 84 citations
- Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation LearningMingcheng Li, Dingkang Yang, Yang Liu, Shunli Wang et al.NeurIPS 2024 · 48 citations
- Transformer-based Feature Reconstruction Network for Robust Multimodal Sentiment AnalysisZiqi Yuan, Wei Li, Hua Xu, Wenmeng YuACM MM 2021 · 186 citations
- Mitigating Inconsistencies in Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Jiantao Zhou, Tianyi LiuEMNLP 2022 · 29 citations
- A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing ModalitiesMingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang et al.AAAI 2024 · 71 citations
