RunawayEvil: Jailbreaking the Image-to-Video Generative Models
yueming lyu, Rufan Qian, Yueming Lyu, Qinglong Liu, Linzhuang Zou, Jie Qin, Songhua Liu, Caifeng Shan
Abstract
Image-to-Video (I2V) generation represents a frontier in content creation, where models synthesize dynamic visual sequences by jointly reasoning from both image and text prompts. This multimodal grounding enables diverse controllability over video attributes. However, it is precisely this capability that introduces a critical security blind spot: by exploiting the interplay between visual and textual cues, attackers can launch multimodal jailbreak attacks that severely compromise output security. Despite the increasing implementation of security mechanisms in real-world I2V systems, such cross-modal threats remain unexplored. Existing attack methods remain confined to single-modal settings, relying solely on isolated text or image perturbations, which severely limits their effectiveness. To bridge this gap, we propose Runaway Evil, the first multimodal jailbreaking framework for I2V models with dynamic evolutionary capability. Built on a Strategy-Tactic-Action paradigm, our framework exhibits self-amplifying attack through three core components: (1) a strategy-aware command unit that enables the attack to self-evolve its strategies through reinforcement learning-driven strategy customization and large language model (LLM)-based strategy exploration; (2) a multimodal tactical planning unit that generates synergistic text jailbreak instructions and image tampering guidelines based on the selected strategies; and (3) an tactical action Unit executes and evaluates the coordinated attacks. This self-evolving architecture allows the framework to continuously adapt and intensify its attack strategies without human intervention. Extensive experiments demonstrate that Runaway Evil achieves state-of-the-art attack success rates on commercial I2V models, such as Open-Sora 2.0 and CogVideoX. This work provides a critical tool for probing and mitigating multimodal vulnerabilities, laying a foundation for building more robust video generation systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Exposing and Defending the Achilles' Heel of Video Mixture-of-ExpertsSongping Wang, Qinglong Liu, Yueming Lyu, Ning Li et al.ICLR 2026 · 3 citations
- Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object DetectionHuafeng Chen, Chenguang Zhu, Yueming Lyu, Caifeng ShanCVPR 2026
Builds on8
- SneakyPrompt: Jailbreaking Text-to-image Generative ModelsYuchen Yang, Bo Hui, Haolin Yuan, Neil Gong et al.S&P 2024 · 188 citations
- Perception-Guided Jailbreak Against Text-to-Image ModelsYihao Huang, Le Liang, Tianlin Li, Xiaojun Jia et al.AAAI 2025 · 34 citations
- UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated ImagesYiting Qu, Xinyue Shen, Yixin Wu, Michael Backes et al.CCS 2025 · 1 citation
- Conditional Image-to-Video Generation with Latent Flow Diffusion ModelsHaomiao Ni, Changhao Shi, Kai Li, Sharon X. Huang et al.CVPR 2023
- Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-to-Image Generation ModelsYingkai Dong, Xiangtao Meng, Ning Yu, Zheng Li et al.S&P 2025
Related papers
- T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak AttacksJiayang Liu, Siyuan Liang, Shiqian Zhao, Rong-Cheng Tu et al.NeurIPS 2025 · 18 citations
- Breaking Multimodal LLM Safety via Video-Driven PromptingDong Wang, XIANGYU HE, Xinqi Lyu, Bin XiaoCVPR 2026
- Reason2Attack: Jailbreaking Text-to-Image Models via LLM ReasoningChenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li et al.AAAI 2026 · 7 citations
- Jailbreak Large Vision-Language Models Through Multi-Modal LinkageYu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang et al.ACL 2025 · 51 citations
- TVChain: Leveraging Textual-Visual Prompt Chains for Jailbreaking Large Vision-Language ModelsHao Yu, Ke Liang, Junxian Duan, Jun Wang et al.AAAI 2026
