Intra- and Inter-Modal Curriculum for Multimodal Learning
Yuwei Zhou, Xin Wang, Hong Chen, Xuguang Duan, Wenwu Zhu
Abstract
Multimodal learning has been widely studied and applied due to its improvement over previous unimodal tasks and its effectiveness on emerging multimodal challenges. However, it has been reported that modal encoders are under-optimized in multimodal learning in contrast to unimodal learning, especially when some modalities are dominant over others. Existing solutions to this problem suffer from two limitations: i) they merely focus on inter-modal balance, failing to consider the influence of intra-modal data on each modality; ii) their implementations heavily rely on unimodal performances or losses, thus being suboptimal for the tasks requiring modal interactions (e.g., visual question answering). To tackle these limitations, we propose I 2 MCL, a generic Intra-and Inter-Modal Curriculum Learning framework which simultaneously considers both data difficulty and modality balance for multimodal learning. In the intra-modal curriculum, we adopt a pretrained teacher model to obtain knowledge distillation loss as the difficulty measurer, which determines the data weights within the corresponding modality. In the inter-modal curriculum, we utilize a Pareto optimization strategy to measure and compare the gradients from distillation loss and task loss across modalities, capable of determining whether a modality should learn from the task or its teacher. Empirical experiments on various tasks including multimodal classification, visual question answering and visual entailment demonstrate that our proposed I 2 MCL is able to tackle the under-optimized modality problem and bring consistent improvement to multimodal learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29ec79c9-9285-4404-98c2-1a6eb15e4137Cited by top-tier papers12
- MotionGPT3: Human Motion as a Second ModalityBingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang et al.ICLR 2026 · 43 citations
- Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal LearningDivyam Madaan, Taro Makino, Sumit Chopra, Kyunghyun ChoNeurIPS 2024 · 19 citations
- OVG-HQ: Online Video Grounding with Hybrid-Modal QueriesRunhao Zeng, Jiaqi Mao, Minghao Lai, Minh Hieu Phan et al.ICCV 2025 · 4 citations
- CMoB: Modality Valuation via Causal Effect for Balanced Multimodal LearningJun Wang, Fuyuan Cao, Zhixin Xue, Xingwang Zhao et al.NeurIPS 2025 · 4 citations
- IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale BenchmarkZhe Cao, Jin Zhang, Ruiheng ZhangICCV 2025 · 2 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural NetworksNan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. GerasICML 2022 · 124 citations
Related papers
- From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringMingrui Lao, Yanming Guo, Yu Liu, Wei Chen et al.ACM MM 2021 · 22 citations
- G2D: Boosting Multimodal Learning with Gradient-Guided DistillationMohammed Rakib, Arunkumar BagavathiICCV 2025 · 1 citation
- Boosting Multimodal Learning via Disentangled Gradient LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 9 citations
- Modality-Balanced Learning for Multimedia RecommendationJinghao Zhang, Guofan Liu, Qiang Liu, Shu Wu et al.ACM MM 2024 · 21 citations
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 86 citations
