Beyond Static Alignment: Adaptive Arbitration for Semantic Incongruence in Semi-Supervised Multimodal Sentiment Analysis
Huicong Li, Xiangbo Ji, Wei Wu
Abstract
Multimodal sentiment analysis is fundamentally challenged by semantic incongruence, where ambiguous visual signals often conflict with explicit textual cues. In semi-supervised scenarios, naively fusing such noisy features contaminates the joint representation, while conventional static alignment strategies fail to effectively arbitrate conflicting modalities in this task, leading to error reinforcement during self-training. To this end, we propose a novel Adaptive Arbitration for Semantic Incongruence (A2SI) framework for semi-supervised multimodal sentiment analysis, which emphasizes stable cross-modal representations and reliable supervision. Specifically, we first constrain unreliable visual representations by leveraging the reliable textual modality as an anchor to align divergent embeddings and reduce representation noise. Based on this, we further consider the reliability of supervision signals and calibrate pseudo-labels by adaptively weighting evidentiary confidence from heterogeneous views. Finally, to prevent error accumulation caused by unreliable samples, we introduce a progressive arbitration mechanism that verifies pseudo-labeled data from dual perspectives, enabling the model to dynamically balance sample diversity and label purity throughout selftraining. Extensive experiments on the MVSA-Single and MVSA-Multiple datasets demonstrate that A2SI consistently outperforms stateof-the-art methods under label-limited settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- Dash: Semi-Supervised Learning with Dynamic ThresholdingYi Xu, Lei Shang, Jinxing Ye, Qi Qian et al.ICML 2021 · 287 citations
Related papers
- SPP-SCL: Semi-Push-Pull Supervised Contrastive Learning for Image-Text Sentiment Analysis and BeyondJiesheng Wu, Shengrong LiAAAI 2026
- Conflict-Aware Adaptive Cross-Reconstruction for Multimodal Sentiment AnalysisYan Wang, Fuyuan Cao, Xingwang ZhaoCVPR 2026
- Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionWuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng YangCVPR 2025
- Pseudo-Label Calibration Semi-supervised Multi-Modal Entity AlignmentLuyao Wang, Pengnian Qi, Xigang Bao, Chunlai Zhou et al.AAAI 2024 · 21 citations
- CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment AnalysisHaoyu Jiang, Xiaoliang Chen, Duoqian Miao, Xiaolin Qin et al.CVPR 2026
