ERL-MR: Harnessing the Power of Euler Feature Representations for Balanced Multi-modal Learning
Weixiang Han, Chengjun Cai, Yu Guo, Jialiang Peng
Abstract
Multi-modal learning leverages data from diverse perceptual media to obtain enriched representations, thereby empowering machine learning models to complete more complex tasks. However, recent research results indicate that multi-modal learning still suffers from " modality imbalance '': Certain modalities' contributions are suppressed by dominant ones, consequently constraining the overall performance enhancement of multimodal learning. To tackle this issue, current approaches attempt to mitigate modality competition in various ways, but their effectiveness is still limited. To this end, we propose an Euler Representation Learning-based Modality Rebalance (ERL-MR) strategy, which reshapes the underlying competitive relationships between modalities into mutually reinforcing win-win situations while maintaining stable feature optimization directions. Specifically, ERL-MR employs Euler's formula to map original features to complex space, constructing cooperatively enhanced non-redundant features for each modality, which helps reverse the situation of modality competition. Moreover, to counteract the performance degradation resulting from optimization drift among modalities, we propose a Multi-Modal Constrained (MMC) loss based on cosine similarity of complex feature phase and cross-entropy loss of individual modalities, guiding the optimization direction of the fusion network. Extensive experiments conducted on four multi-modal multimedia datasets and two task-specific multi-modal multimedia datasets demonstrate the superiority of our ERL-MR strategy over state-of-the-art baselines, achieving modality rebalancing and further performance improvements.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c6067995-bff4-4b3d-9b3d-dac9f2de0a88Cited by top-tier papers1
Ask how each one uses itRelated papers
- PMR: Prototypical Modal Rebalance for Multimodal LearningYunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang et al.CVPR 2023
- Improving Multimodal Learning via Imbalanced LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 7 citations
- Asymmetric Reinforcing Against Multi-Modal Representation BiasXiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang et al.AAAI 2025 · 6 citations
- Balancing Multimodal Training Through Game-Theoretic RegularizationKonstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang et al.NeurIPS 2025 · 17 citations
- Uncertainty-Guided Modal Rebalance for Hateful Memes DetectionChuanpeng Yang, Yaxin Liu, Fuqing Zhu, Jizhong Han et al.ACL 2024
