New Wide-Net-Casting Jailbreak Attacks Risk Large Models
Qiuchi Xiang, Haoxuan Qu, Hossein Rahmani, Jun Liu
摘要
Jailbreak attacks on large models have drawn growing attention due to their close ties to societal safety. This work identifies a practical yet unexplored jailbreak scenario, the wide-net-casting scenario, where an adversary can query a group of large models instead of a single one to elicit harmful outputs. Our analysis reveals substantial yet previously overlooked safety risks under this scenario. As a key part of our analysis, we further develop a novel jailbreak method tailored to the wide-net-casting scenario. With this tailored method, the jailbreak success rate can even reach 100% in some experiments when targeting the large models without additional safeguards, exposing wide-net-casting as a distinct, high-risk scenario that warrants attention in future evaluation and defense research. Code is available here. Warning: This paper contains potentially harmful example text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao 等ICLR 2024 · 被引用 831 次
相关 Paper
- Attention Eclipse: Manipulating Attention to Bypass LLM Safety-AlignmentPedram Zaree, Md Abdullah Al Mamun, Quazi Mishkatul Alam, Yue Dong 等EMNLP 2025
- Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language ModelsLang Gao, Jiahui Geng, Xiangliang Zhang, Preslav Nakov 等ACL 2025
- From LLMs to MLLMs: Exploring the Landscape of Multimodal JailbreakingSiyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu WeiEMNLP 2024 · 被引用 5 次
- Weak-to-Strong Jailbreaking on Large Language ModelsXuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du 等ICML 2025
- SoK: Robustness in Large Language Models against Jailbreak AttacksFeiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang 等S&P 2026 · 被引用 4 次
