The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model Services
Jian Cui, Mingming Zha, XiaoFeng Wang, Xiaojing Liao
摘要
Web content is an essential element for large language model (LLM) services, supporting both training and inference processes. To manage the content access of web bots from LLM service vendors (i.e., LLM bots), web content publishers are increasingly incorporated content access rules into robots.txt, a long-established web content management protocol. However, the rise of proprietary LLM bots, such as OpenAI's ChatGPT-User and Google's Google-Extended, has raised concerns about the transparency of web content access and whether these bots adherence to robots.txt rules. However, there is limited understanding of these LLM bots, concerning their impact on web publishers and broader web content governance. To fill this gap, we present a systematic analysis of 18 LLM bots on 582,281 robots.txt files. Our findings reveal a significant increase in robots.txt rules associated with LLM bots, particularly in domains that fall into the finance and news category. Despite the heightened integration, web publishers face challenges in managing robots.txt configurations due to the complexity of the LLM ecosystem and the involvement of third-party brokers. Furthermore, we identified several cases of robots.txt violations, including instances where LLMs memorized web content from restricted domains, and where ChatGPT-User ignored robots.txt and accessed restricted content. These results highlight the gaps in the current web content governance and underscore the need for enforceable content management mechanisms to respect web publishers' intentions and content control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Identifying AI Web Scrapers Using Canary TokensSteven Seiden, Triss Ren, Caroline Zhang, Taein Kim 等CCS 2026
- A Large-scale Measurement of In-Page Prompt Injections Against LLM Web AgentsSoheil Khodayari, Xuenan Zhang, Bhupendra Acharya, Giancarlo PellegrinoCCS 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 被引用 304 次
相关 Paper
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the WebNicolas Steinacker-Olsztyn, Devashish Gosain, Ha DaoWWW 2026
- MDEval: Evaluating and Enhancing Markdown Awareness in Large Language ModelsZhongpu Chen, Yinfeng Liu, Long Shi, Zhi-Jie Wang 等WWW 2025 · 被引用 12 次
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang 等NDSS 2024
- From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language ModelsSayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, Shirin NilizadehS&P 2024 · 被引用 57 次
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot PluginsYigitcan Kaya, Anton Landerer, Stijn Pletinckx, Michelle Zimmermann 等S&P 2026 · 被引用 12 次
