Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal Learning
Hossein Rajoli Nowdeh, Jie Ji, Xiaolong Ma, Fatemeh Afghah
摘要
In multimodal learning, dominant modalities often overshadow others, limiting generalization. We propose Modality-Aware Sharpness-Aware Minimization (M-SAM), a model-agnostic framework that applies to many modalities and supports early and late fusion scenarios. In every iteration, M-SAM in three steps optimizes learning. First, it identifies the dominant modality based on modalities' contribution in the accuracy using Shapley. Second, it decomposes the loss landscape, or in another language, it modulates the loss to prioritize the robustness of the model in favor of the dominant modality, and third, M-SAM updates the weights by backpropagation of modulated gradients. This ensures robust learning for the dominant modality while enhancing contributions from others, allowing the model to explore and exploit complementary features that strengthen overall performance. Extensive experiments on four diverse datasets show that M-SAM outperforms the latest state-of-the-art optimization and gradient manipulation methods and significantly balances and improves multimodal learning. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 被引用 388 次
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang 等CVPR 2022 · 被引用 264 次
- Surrogate Gap Minimization Improves Sharpness-Aware TrainingJuntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui 等ICLR 2022 · 被引用 213 次
相关 Paper
- MASAM: Multimodal Adaptive Sharpness-Aware Minimization for Heterogeneous Data FusionZijie Chen, Kejing Yin, Wenfang Yao, William Kwok-Wai Cheung 等ICLR 2026
- CMoB: Modality Valuation via Causal Effect for Balanced Multimodal LearningJun Wang, Fuyuan Cao, Zhixin Xue, Xingwang Zhao 等NeurIPS 2025 · 被引用 4 次
- Towards Balanced Active Learning for Multimodal ClassificationMeng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou 等ACM MM 2023 · 被引用 5 次
- PaSE: Prototype-aligned Calibration and Shapley-based Equilibrium for Multimodal Sentiment AnalysisKang He, Boyu Chen, Yuzhe Ding, Fei Li 等AAAI 2026 · 被引用 1 次
- Multimodal Representation Learning by Alternating Unimodal AdaptationXiaohui Zhang, Jaehong Yoon, Mohit Bansal, Huaxiu YaoCVPR 2024
