Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration
Chunlei Meng, Pengbin Feng, Rong Fu, Hoi Leong Lee, Xiaojing Du, Yuying Li, Zeyu Zhang, Weilin Zhou, Chun Ouyang, Zhongxue Gan
Abstract
Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers from two limitations: modality dominance, where optimization gravitates towards the path of least resistance, ignoring weaker but informative modalities, and spurious modality coupling, where models overfit to incidental cross-modal correlations. To address these, we propose Group Cognition Learning (GCL), a governed collaboration paradigm that applies a two-stage protocol after modality-specific encoding. In Stage 1 (Selective Interaction), a Routing Agent proposes directed interaction routes, and an Auditing Agent assigns sample-wise gates to emphasize exchanges that yield positive marginal predictive gain while suppressing redundant coupling. In Stage 2 (Consensus Formation), a Public-Factor Agent maintains an explicit shared factor, and an Aggregation Agent produces the final prediction through contribution-aware weighting while keeping each modality representation as a specialization channel. Extensive experiments on CMU-MOSI, CMU-MOSEI, and MIntRec demonstrate that GCL mitigates dominance and coupling, establishing state-of-the-art results across both regression and classification benchmarks. Analysis experiments further demonstrate the effectiveness of the design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment AnalysisJiuding Yang, Yakun Yu, Di Niu, Weidong Guo et al.ACL 2023 · 135 citations
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou et al.ACM MM 2022 · 66 citations
Related papers
- Multimodal Representation Learning by Alternating Unimodal AdaptationXiaohui Zhang, Jaehong Yoon, Mohit Bansal, Huaxiu YaoCVPR 2024
- Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal LearningHossein Rajoli Nowdeh, Jie Ji, Xiaolong Ma, Fatemeh AfghahNeurIPS 2025 · 3 citations
- Closing the Modality Gap Aligns Group-Wise SemanticsEleonora Grassucci, Giordano Cicchetti, Emanuele Frasca, Aurelio Uncini et al.ICLR 2026 · 5 citations
- PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation NetworkKehan Wang, Huan Zhao, Yong Wei, Xupeng Zha et al.AAAI 2026
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian et al.CVPR 2026 · 9 citations
