WebCloak: Characterizing and Mitigating Threats From LLM-Driven Web Agents as Intelligent Scrapers
Xinfeng Li, Tianze Qiu, Yingbin Jin, Lixu Wang, Hanqing Guo, Xiaojun Jia, Xiaofeng Wang, Wei Dong
摘要
The rise of web agents powered by large language models (LLMs) is reshaping the landscape of human-computer interaction, enabling users to automate complex web tasks with natural language commands. However, this progress introduces significant, yet largely unexplored security concerns: adversaries can employ such web agents to conduct largescale web scraping, particularly of visual content. This paper presents the first systematic characterization of the danger represented by such LLM-driven web agents as intelligent scrapers. We develop LLMCrawlBench, a large test set of 237 extracted real-world webpages (10,895 images) from 50 popular high-traffic websites in 5 critical categories, designed specifically for adversarial image extraction evaluation. Our metrics across over 32 scraper implementations, including LLMto-Script (L2S), LLM-Native crawlers (LNC), and LLM-based web agents (LWA), demonstrate that while some tools exhibit working issues, advanced LLM-powered frameworks lower the bar for effective scraping.
Such new agent-as-attacker threats motivate us to introduce WebCloak, an effective, lightweight defense that specifically targets the main weakness of LLM crawler agents' fundamental "Parse-then-Interpret" mechanism. Our key idea is dual-layered: (1) Dynamic Structural Obfuscation, which not only randomizes structural cues but also restores visual content client-side using non-traditional methods less amenable to direct LLM exploitation, and (2) Optimized Semantic Labyrinth to mislead the central LLM interpretation of the agent through added harmless-yetmisleading contextual clues, all while not sacrificing visual quality for legitimate users. Our evaluations demonstrate that WebCloak significantly reduces scraping recall rates from 88.7% to 0% against leading LLM-driven scraping agents, offering a robust and practical countermeasure.
Artist Images News Events Stock Data Travel Data
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu 等ACL 2024 · 被引用 30 次
相关 Paper
- AutoScraper: A Progressive Understanding Web Agent for Web Scraper GenerationWenhao Huang, Zhouhong Gu, Chenghao Peng, Jiaqing Liang 等EMNLP 2024 · 被引用 6 次
- AgentBreaker: Evaluating Context-Aware Indirect Prompt Injection Risks in Modern Web AgentsYongbi Son, Changoo Lee, Dongwon Shin, Byoungyoung Lee 等ISSTA 2026
- Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO ManipulationPei Chen, Geng Hong, Xinyi Wu, Mengying Wu 等WWW 2026
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr 等ICML 2025
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
