Enhancing Adversarial Robustness of Multi-modal Recommendation via Modality Balancing
Yu Shang, Chen Gao, Jiansheng Chen, Depeng Jin, Huimin Ma, Yong Li
Abstract
Recently multi-modal recommender systems have been widely applied in real scenarios such as e-commerce businesses. Existing multi-modal recommendation methods exploit the multi-modal content of items as auxiliary information and fuse them to boost performance. Despite the superior performance achieved by multi-modal recommendation models, there's currently no understanding of their robustness to adversarial attacks. In this work, we first identify the vulnerability of existing multi-modal recommendation models. Next, we show the key reason for such vulnerability is modality imbalance, i.e., the prediction score margin between positive and negative samples in the sensitive modality will drop dramatically facing adversarial attacks and fail to be compensated by other modalities. Finally, based on this finding we propose a novel defense method to enhance the robustness of multi-modal recommendation models through modality balancing. Specifically, we first adopt an embedding distillation to obtain a pair of content-similar but prediction-different item embeddings in the sensitive modality and calculate the score margin reflecting the modality vulnerability. Then we optimize the model to utilize the score margin between positive and negative samples in other modalities to compensate for the vulnerability. The proposed method can serve as a plug-and-play module and is flexible to be applied to a wide range of multi-modal recommendation models. Extensive experiments on two real-world datasets demonstrate that our method significantly improves the robustness of multi-modal recommendation models with nearly no performance degradation on clean data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 227f6267-446e-4042-990a-9c9ca83eadbcCited by top-tier papers5
- Step Vulnerability Guided Mean Fluctuation Adversarial Attack against Conditional Diffusion ModelsHongwei Yu, Jiansheng Chen, Xinlong Ding, Yudong Zhang et al.AAAI 2024 · 18 citations
- VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement LearningSiran Chen, Boyu Chen, Yuxiao Luo, Chenyun Yu et al.AAAI 2026
- From Zero to Hero: Cross-modal-enhanced Adversarial Item Promotion Attack against Multimodal Recommender SystemsMengyu Yao, Ziqi Zhang, Yifeng Cai, Junlin Liu et al.USENIX Security 2026
- Sign-Aware Multimodal Graph RecommendationYahong Lian, Haotian Tian, Chunyao Song, Tingjian GeAAAI 2026
- The Hidden Risk: Membership Inference Attacks on Multimodal Federated Learning via Modality ImbalanceChang Ma, Jun Li, Kang Wei, Yipeng Zhou et al.ICML 2026
Builds on13
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He et al.ACM MM 2020 · 374 citations
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu et al.ACM MM 2021 · 350 citations
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng et al.WWW 2023 · 326 citations
- How Dataset Characteristics Affect the Robustness of Collaborative Recommendation ModelsYashar Deldjoo, Tommaso Di Noia, Eugenio Di Sciascio, Felice Antonio MerraSIGIR 2020 · 50 citations
Related papers
- Modality-Balanced Learning for Multimedia RecommendationJinghao Zhang, Guofan Liu, Qiang Liu, Shu Wu et al.ACM MM 2024 · 21 citations
- Aligning Distillation For Cold-start Item RecommendationFeiran Huang, Zefan Wang, Xiao Huang, Yufeng Qian et al.SIGIR 2023 · 100 citations
- Online Distillation-enhanced Multi-modal Transformer for Sequential RecommendationWei Ji, Xiangyan Liu, An Zhang, Yinwei Wei et al.ACM MM 2023 · 32 citations
- VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender SystemsGuowei Guan, Yurong Hao, Jiaming Zhang, Tiantong Wu et al.ICML 2026
- DVIB: Towards Robust Multimodal Recommender Systems via Variational Information Bottleneck DistillationWenkuan Zhao, Shanshan Zhong, Yifan Liu, Wushao Wen et al.WWW 2025 · 7 citations
