MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning
Seonghyeon Hwang, Soyoung Choi, Steven Euijong Whang
Abstract
Multimodal models often over-rely on dominant modalities, failing to achieve optimal performance. While prior work focuses on modifying training objectives or optimization procedures, data-centric solutions remain underexplored. We propose MIDAS, a novel data augmentation strategy that generates misaligned samples with semantically inconsistent cross-modal information, labeled using unimodal confidence scores to compel learning from contradictory signals. However, this confidence-based labeling can still favor the more confident modality. To address this within our misaligned samples, we introduce weak-modality weighting, which dynamically increases the loss weight of the least confident modality, thereby helping the model fully utilize weaker modality. Furthermore, when misaligned features exhibit greater similarity to the aligned features, these misaligned samples pose a greater challenge, thereby enabling the model to better distinguish between classes. To leverage this, we propose hard-sample weighting, which prioritizes such semantically ambiguous misaligned samples. Experiments on multiple multimodal classification benchmarks demonstrate that MIDAS significantly outperforms related baselines in addressing modality imbalance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Hard Negative Mixing for Contrastive LearningYannis Kalantidis, Mert Bülent Sariyildiz, Noé Pion, Philippe Weinzaepfel et al.NeurIPS 2020 · 805 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
Related papers
- Towards Balanced Active Learning for Multimodal ClassificationMeng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou et al.ACM MM 2023 · 5 citations
- MASAM: Multimodal Adaptive Sharpness-Aware Minimization for Heterogeneous Data FusionZijie Chen, Kejing Yin, Wenfang Yao, William Kwok-Wai Cheung et al.ICLR 2026
- Asymmetric Reinforcing Against Multi-Modal Representation BiasXiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang et al.AAAI 2025 · 6 citations
- Mask to Align, Weight to Disambiguate: Reliable Unsupervised Cross-Modal Hashing with Masked-Weight ContrastFan Yang, Yuanzhi Zhao, Haimei Zhao, Yudong Zhao et al.CVPR 2026
- Adaptive Re-calibration Learning for Balanced Multimodal Intention RecognitionQu Yang, Xiyang Li, Fu Lin, Mang YeNeurIPS 2025 · 2 citations
