PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network
Kehan Wang, Huan Zhao, Yong Wei, Xupeng Zha, Guanghui Ye, Cheng Zhu, Yiming Liu, Zixing Zhang
Abstract
Existing multimodal representation learning approaches often rely on simple feature concatenation or unified transformations, which fail to effectively disentangle and leverage common and private information across different modalities in a progressive manner. Moreover, they typically lack adaptive modeling tailored to specific task requirements. To address these limitations, we propose a Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network (PLUM-Net). It first employs a multilevel semantic alignment module to synchronize global and local semantics across audio, visual and textual streams. On this aligned foundation, a prototype-based single-modal label generation module derives modality-specific hard and soft-labels that subtly steer the network toward a cleaner split between shared and private cues. Guided by these labels, the task-conditioned feature bifurcator module channels information through the most beneficial common or private pathway for the given task, after which a private refinement module polishes and fuses each modality’s idiosyncratic signals. Extensive experiments show that PLUM-Net delivers strong performance on datasets such as CMU-MOSI, CMU-MOSEI and UR-FUNNY, achieving an ACC-2 of 90.3% on CMU-MOSI, representing a 2%–4% improvement over previous SOTA models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eeac0fc0-b25c-4053-95c2-d11dd932076bBuilds on16
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionGuimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu et al.EMNLP 2022 · 206 citations
- CM-BERT: Cross-Modal BERT for Text-Audio Sentiment AnalysisKaicheng Yang, Hua Xu, Kai GaoACM MM 2020 · 129 citations
Related papers
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
- Tailor Versatile Multi-Modal Learning for Multi-Label Emotion RecognitionYi Zhang, Mingyuan Chen, Jundong Shen, Chongjun WangAAAI 2022 · 92 citations
- Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal AlignmentLinlin Ji, Li LiuAAAI 2026
- Cross-modality Representation Interactive Learning for Multimodal Sentiment AnalysisJian Huang, Yanli Ji, Yang Yang, Heng Tao ShenACM MM 2023 · 17 citations
- DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisPan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen et al.AAAI 2025 · 84 citations
