URLcoat: Exploiting Web Search Capability to Jailbreak Large Language Models
Yiheng Sun, Linkang Du, Zhou Su, Yuntao Wang, Han Liu
Abstract
Large language models (LLMs) achieve remarkable advances in understanding and reasoning with human language, which are widely applied in software development, content creation, healthcare, etc. However, their vulnerabilities to security threats, especially jailbreak attacks, remain a significant issue. Existing research on jailbreak mainly focuses on the security risks of LLMs' inherent thinking and reasoning, overlooking the new attack surface introduced by web search capability. Attackers can exploit this by guiding LLMs to retrieve information from external URLs, which is then used to implicitly reconstruct harmful instructions and circumvent safety mechanisms, leading to the generation of harmful content. In this paper, we propose a novel jailbreak attack, named URLcoat. The core idea is to exploit the web search capabilities of LLMs to circumvent their security safeguards. URLcoat incorporates three core strategies: obfuscating the feature of sensitive words to evade input detection, reconstructing harmful instructions via implicit associations with external URLs, and contextual narrative guidance to bypass output filtering. The experimental results reveal that URLcoat attains 100 % attack success rates in mainstream LLMs, including GPT5, GPT-4o, Gemini 2.0 Flash Thinking, DeepSeek R1, Kimi 1.5, ChatGLM-4, Grok 3, Gemini 2.5 Pro, Gemini 2.5 Flash, and Gemini 2.0 Flash, exceeding the performance of state-of-the-art jailbreak techniques. This study examines the security vulnerabilities arising from lLMs' web search capability, which facilitates the LLMs to produce harmful output.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 736f6eba-409c-4762-8f0b-8c719fcc2d93Related papers
- MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative StrategiesWeiwei Qi, Shuo Shao, Wei Gu, Tianhang Zheng et al.AAAI 2026
- h4rm3l: A Language for Composable Jailbreak Attack SynthesisMoussa Koulako Bala Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi et al.ICLR 2025
- DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image InputsWenzhuo Xu, Zhipeng Wei, Zonghao Ying, Deyue Zhang et al.ACL 2026
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang et al.NDSS 2024
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu et al.AAAI 2025 · 26 citations
