Identifying AI Web Scrapers Using Canary Tokens
Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger
摘要
From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by AI chatbots such as ChatGPT and Claude. However, large-scale web scraping to feed such chatbots can affect site stability and raise legal, privacy, or ethics concerns. If website owners wish to limit AI-chatbot-related web scraping on their site, due to these or other concerns, they may turn to scraper access control mechanisms like the Robots Exclusion Protocol. To be most effective, such mechanisms require site owners to first identify the scrapers that they wish to restrict (e.g., via User-Agent strings). Existing mechanisms to identify AI-chatbot-related scrapers rely on voluntary disclosure by companies, one-off research experiments, or crowd-sourced reports -none of which are reliable or scalable.
This paper proposes a novel technique for accurately and automatically inferring AI-chatbot-related scrapers. We host dynamic websites that serve unique canary tokens to each visiting scraper, then prompt AI chatbots for information about our sites. If a chatbot consistently generates outputs containing tokens unique to a scraper, it provides evidence of exposure to that scraper. Via experiments across 18 production chatbots, we demonstrate that our approach can reliably identify which scrapers feed which chatbot, including several that are not publicly known or disclosed by the companies. Our approach provides a promising avenue for unprivileged third parties to infer which scrapers serve data to which chatbots, enabling better control over and transparency into unwanted scraping.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative ModelsShawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu 等S&P 2024 · 被引用 102 次
- Good Bot, Bad Bot: Characterizing Automated Browsing ActivityXigao Li, Babak Amin Azad, Amir Rahmati, Nick NikiforakisS&P 2021 · 被引用 45 次
- Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language ModelsYisheng Zhong, Yizhu Wen, Junfeng Guo, Mehran Kafai 等EMNLP 2025
- Glaze: Protecting Artists from Style Mimicry by Text-to-Image ModelsShawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng 等USENIX Security 2023
相关 Paper
- The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model ServicesJian Cui, Mingming Zha, XiaoFeng Wang, Xiaojing LiaoCCS 2025
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the WebNicolas Steinacker-Olsztyn, Devashish Gosain, Ha DaoWWW 2026
- AutoScraper: A Progressive Understanding Web Agent for Web Scraper GenerationWenhao Huang, Zhouhong Gu, Chenghao Peng, Jiaqing Liang 等EMNLP 2024 · 被引用 6 次
- DEMASQ: Unmasking the ChatGPT WordsmithKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiNDSS 2024
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot PluginsYigitcan Kaya, Anton Landerer, Stijn Pletinckx, Michelle Zimmermann 等S&P 2026 · 被引用 12 次
