Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models
Yisheng Zhong, Yizhu Wen, Junfeng Guo, Mehran Kafai, Heng Huang, Hanqing Guo, Zhuangdi Zhu
Abstract
The protection of cyber Intellectual Property (IP) such as web content is an increasingly critical concern. The rise of large language models (LLMs) with online retrieval capabilities enables convenient access to information but often undermines the rights of original content creators. As users increasingly rely on LLM-generated responses, they gradually diminish direct engagement with original information sources, which will significantly reduce the incentives for IP creators to contribute, and lead to a saturating cyberspace with more AIgenerated content. In response, we propose a novel defense framework that empowers web content creators to safeguard their web-based IP from unauthorized LLM real-time extraction and redistribution by leveraging the semantic understanding capability of LLMs themselves. Our method follows principled motivations and effectively addresses an intractable black-box optimization problem. Real-world experiments demonstrated that our methods improve defense success rates from 2.5% to 88.6% on different LLMs, outperforming traditional defenses such as configuration-based restrictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- WebCloak: Characterizing and Mitigating Threats From LLM-Driven Web Agents as Intelligent ScrapersXinfeng Li, Tianze Qiu, Yingbin Jin, Lixu Wang et al.S&P 2026 · 13 citations
- Confundo: Learning to Generate Robust Poison for Practical RAG SystemsHaoyang Hu, Zhejun Jiang, Yueming Lyu, Junyuan Zhang et al.USENIX Security 2026 · 5 citations
- DUET: Distilled LLM Unlearning from an Efficiently Contextualized TeacherYisheng Zhong, Zhengbang Yang, Zhuangdi ZhuICLR 2026 · 4 citations
- Identifying AI Web Scrapers Using Canary TokensSteven Seiden, Triss Ren, Caroline Zhang, Taein Kim et al.CCS 2026
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman et al.KDD 2025 · 27 citations
Related papers
- REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language ModelsRuisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, Farinaz KoushanfarUSENIX Security 2024 · 88 citations
- Imperceptible Content Poisoning in LLM-Powered ApplicationsQuan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng et al.ASE 2024 · 3 citations
- SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text GenerationXiaoze Liu, Ting Sun, Tianyang Xu, Feijie Wu et al.EMNLP 2024 · 4 citations
- Prompt Obfuscation for Large Language ModelsDavid Pape, Sina Mavali, Thorsten Eisenhofer, Lea SchönherrUSENIX Security 2025
- Protecting Intellectual Property of Large Language Model-Based Code Generation APIs via WatermarksZongjie Li, Chaozheng Wang, Shuai Wang, Cuiyun GaoCCS 2023 · 25 citations
