USENIX Security2026Top-tier venue
From Zero to Hero: Cross-modal-enhanced Adversarial Item Promotion Attack against Multimodal Recommender Systems
Mengyu Yao, Ziqi Zhang, Yifeng Cai, Junlin Liu, Xinyi Fu, Weiqiang Wang, Xiangqun Chen, Yao Guo, Ding Li
Abstract
Multimodal recommender systems (MRSs) jointly leverage visual and textual item representations to determine item ranking and exposure in modern online platforms. Their strong dependence on item content, however, introduces new security risks. In particular, malicious sellers can conduct the adversarial item promotion (AIP) attack to manipulate item content to deceive the recommender into overranking specific items. When it succeeds, the promoted items gain disproportionately higher exposure, leading to significant market visibility and direct economic gain. However, existing AIP research primarily targets unimodal recommenders, leaving the unique vulnerabilities of MRSs largely unexplored. Meanwhile, multimodal adversarial attacks for vision–language models (VLMs) optimize objectives unrelated to ranking and lack cross-modal coordination unique to MRSs. To bridge this gap, we propose CREAM (CRoss-modal-Enhanced AIP attack against MRSs). Our key insight is to jointly perturb multiple modalities in a semantically consistent manner. We integrate a tailored visual perturbator, a text generator, and a joint optimization controller to fully exploit cross-modal correlations in a black-box setting. Our comprehensive evaluation shows that CREAM significantly outperforms existing methods, achieving on average 5.75x and 2.89x higher exposure at top-10 and top-50 metrics, and demonstrates robustness under evolving real-world conditions. This exposes a tangible economic risk to recommender platforms. Meanwhile, CREAM maintains high imperceptibility across visual, textual, and cross-modal dimensions. We further investigate several potential defense strategies and demonstrate their limitations, highlighting the urgent need for stronger protections against adversarial threats in MRSs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
Related papers
- Adversarial Item Promotion: Vulnerabilities at the Core of Top-N Recommenders that Use Images to Address Cold StartZhuoran Liu, Martha A. LarsonWWW 2021 · 34 citations
- VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender SystemsGuowei Guan, Yurong Hao, Jiaming Zhang, Tiantong Wu et al.ICML 2026
- Stealthy Attack on Large Language Model based RecommendationJinghao Zhang, Yuting Liu, Qiang Liu, Shu Wu et al.ACL 2024
- Prompt-Unknown Promotion Attacks against LLM-based Sequential Recommender SystemsYuchuan Zhao, Tong Chen, Junliang Yu, Zongwei Wang et al.SIGIR 2026
- Enhancing Adversarial Robustness of Multi-modal Recommendation via Modality BalancingYu Shang, Chen Gao, Jiansheng Chen, Depeng Jin et al.ACM MM 2023 · 9 citations
