Identifying AI Web Scrapers Using Canary Tokens
Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger
Abstract
From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by AI chatbots such as ChatGPT and Claude. However, large-scale web scraping to feed such chatbots can affect site stability and raise legal, privacy, or ethics concerns. If website owners wish to limit AI-chatbot-related web scraping on their site, due to these or other concerns, they may turn to scraper access control mechanisms like the Robots Exclusion Protocol. To be most effective, such mechanisms require site owners to first identify the scrapers that they wish to restrict (e.g., via User-Agent strings). Existing mechanisms to identify AI-chatbot-related scrapers rely on voluntary disclosure by companies, one-off research experiments, or crowd-sourced reports -none of which are reliable or scalable.
This paper proposes a novel technique for accurately and automatically inferring AI-chatbot-related scrapers. We host dynamic websites that serve unique canary tokens to each visiting scraper, then prompt AI chatbots for information about our sites. If a chatbot consistently generates outputs containing tokens unique to a scraper, it provides evidence of exposure to that scraper. Via experiments across 18 production chatbots, we demonstrate that our approach can reliably identify which scrapers feed which chatbot, including several that are not publicly known or disclosed by the companies. Our approach provides a promising avenue for unprivileged third parties to infer which scrapers serve data to which chatbots, enabling better control over and transparency into unwanted scraping.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative ModelsShawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu et al.S&P 2024 · 102 citations
- Good Bot, Bad Bot: Characterizing Automated Browsing ActivityXigao Li, Babak Amin Azad, Amir Rahmati, Nick NikiforakisS&P 2021 · 45 citations
- Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language ModelsYisheng Zhong, Yizhu Wen, Junfeng Guo, Mehran Kafai et al.EMNLP 2025
- Glaze: Protecting Artists from Style Mimicry by Text-to-Image ModelsShawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng et al.USENIX Security 2023
Related papers
- The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model ServicesJian Cui, Mingming Zha, XiaoFeng Wang, Xiaojing LiaoCCS 2025
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the WebNicolas Steinacker-Olsztyn, Devashish Gosain, Ha DaoWWW 2026
- AutoScraper: A Progressive Understanding Web Agent for Web Scraper GenerationWenhao Huang, Zhouhong Gu, Chenghao Peng, Jiaqing Liang et al.EMNLP 2024 · 6 citations
- DEMASQ: Unmasking the ChatGPT WordsmithKavita Kumari, Alessandro Pegoraro, Hossein Fereidooni, Ahmad-Reza SadeghiNDSS 2024
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot PluginsYigitcan Kaya, Anton Landerer, Stijn Pletinckx, Michelle Zimmermann et al.S&P 2026 · 12 citations
