Lune

CCS2026顶会

Identifying AI Web Scrapers Using Canary Tokens

Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger

2026年份

摘要

From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by AI chatbots such as ChatGPT and Claude. However, large-scale web scraping to feed such chatbots can affect site stability and raise legal, privacy, or ethics concerns. If website owners wish to limit AI-chatbot-related web scraping on their site, due to these or other concerns, they may turn to scraper access control mechanisms like the Robots Exclusion Protocol. To be most effective, such mechanisms require site owners to first identify the scrapers that they wish to restrict (e.g., via User-Agent strings). Existing mechanisms to identify AI-chatbot-related scrapers rely on voluntary disclosure by companies, one-off research experiments, or crowd-sourced reports -none of which are reliable or scalable.

This paper proposes a novel technique for accurately and automatically inferring AI-chatbot-related scrapers. We host dynamic websites that serve unique canary tokens to each visiting scraper, then prompt AI chatbots for information about our sites. If a chatbot consistently generates outputs containing tokens unique to a scraper, it provides evidence of exposure to that scraper. Via experiments across 18 production chatbots, we demonstrate that our approach can reliably identify which scrapers feed which chatbot, including several that are not publicly known or disclosed by the companies. Our approach provides a promising avenue for unprivileged third parties to infer which scrapers serve data to which chatbots, enabling better control over and transparency into unwanted scraping.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖