Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
Changyue Li, Jiaying Li, Youliang Yuan, Jiaming He, Zhicong Huang, Pinjia He
Abstract
Multimodal Large Language Models (MLLMs) are increasingly deployed in stateless systems, such as autonomous driving and robotics. This paper investigates a novel threat: Semantic-Aware Hijacking. We explore the feasibility of hijacking multiple stateless decisions simultaneously using a single universal perturbation. We introduce the Semantic-Aware Universal Perturbation (SAUP), which acts as a semantic router, "actively" perceiving input semantics and routing them to distinct, attacker-defined targets. To achieve this, we conduct a theoretical and empirical analysis on the geometric properties in the latent space. Guided by these insights, we propose the Semantic-Oriented (SORT) optimization strategy and annotate a new dataset with fine-grained semantics to evaluate performance. Extensive experiments on three representative MLLMs demonstrate the fundamental feasibility of this attack, achieving a 66% attack success rate over five targets using a single frame against Qwen. MLLM A Noise Frame Paste Drive as I told you.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17f2ca66-7440-4b5c-8441-a61dc2166cf5Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan Shayegani, Yue Dong, Nael B. Abu-GhazalehICLR 2024 · 271 citations
Related papers
- Phi: Preference Hijacking in Multi-modal Large Language Models at Inference TimeYifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin et al.EMNLP 2025
- Efficient Universal Goal Hijacking with Semantics-guided Prompt OrganizationYihao Huang, Chong Wang, Xiaojun Jia, Qing Guo et al.ACL 2025
- Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt InjectionMeng Chen, Kun Wang, Li Lu, Jiaheng Zhang et al.S&P 2026 · 3 citations
- LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained ModelsAlvi Md. Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Chris ThomasAAAI 2026
- Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language ModelsYanyun Wang, Yu Huang, Zi Liang, Xixin Wu et al.ICML 2026
