Balancing Multimodal Training Through Game-Theoretic Regularization
Konstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang, Matthew B. Blaschko, Maarten De Vos
Abstract
Multimodal learning holds promise for richer information extraction by capturing dependencies across data sources. Yet, current training methods often underperform due to modality competition, a phenomenon where modalities contend for training resources leaving some underoptimized. This raises a pivotal question: how can we address training imbalances, ensure adequate optimization across all modalities, and achieve consistent performance improvements as we transition from unimodal to multimodal data? This paper proposes the Multimodal Competition Regularizer (MCR), inspired by a mutual information (MI) decomposition designed to prevent the adverse effects of competition in multimodal training. Our key contributions are: 1) A game-theoretic framework that adaptively balances modality contributions by encouraging each to maximize its informative role in the final prediction 2) Refining lower and upper bounds for each MI term to enhance the extraction of both task-relevant unique and shared information across modalities. 3) Proposing latent space permutations for conditional MI estimation, significantly improving computational efficiency. MCR outperforms all previously suggested training strategies and simple baseline, clearly demonstrating that training modalities jointly leads to important performance gains on both synthetic and large real-world datasets. We release our code and models at https://github.com/kkontras/MCR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b0c5919-890d-450f-8d3d-7a7a21ec3bdaCited by top-tier papers3
- Rethinking Multimodal Learning from the Perspective of Mitigating Classification Ability DisproportionQing-Yuan Jiang, Longfei Huang, Yang YangNeurIPS 2025 · 15 citations
- THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENTStefanos Koutoupis, Michaela Areti Zervou, Konstantinos Kontras, Maarten De Vos et al.CVPR 2026 · 5 citations
- Information-Theoretic Decomposition for Multimodal Interaction LearningZequn Yang, Yake Wei, Haotian Ni, Zhihao Xu et al.CVPR 2026 · 1 citation
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Multimodal Classification via Total Correlation MaximizationFeng Yu, Xiangyu Wu, Yang Yang, Jianfeng LuICLR 2026 · 4 citations
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsLiangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng et al.NeurIPS 2025 · 10 citations
- Improving Multimodal Learning via Imbalanced LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 7 citations
- ERL-MR: Harnessing the Power of Euler Feature Representations for Balanced Multi-modal LearningWeixiang Han, Chengjun Cai, Yu Guo, Jialiang PengACM MM 2024
- Learning Optimal Multimodal Information Bottleneck RepresentationsQilong Wu, Yiyang Shao, Jun Wang, Xiaobo SunICML 2025
