Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web
Nicolas Steinacker-Olsztyn, Devashish Gosain, Ha Dao
Abstract
Large Language Models (LLMs) are increasingly relying on web crawling to stay up to date and accurately answer user queries. These crawlers are expected to honor robots[.]txt files, which govern automated access. In this study, for the first time, we investigate whether reputable news websites and misinformation sites differ in how they configure these files, particularly in relation to AI crawlers. Analyzing a curated dataset, we find a stark contrast: 60.0% of reputable sites disallow at least one AI crawler, compared to just 9.1% of misinformation sites in their robots[.]txt files. Reputable sites forbid an average of 15.5 AI user agents, while misinformation sites prohibit fewer than one. We then measure active blocking behavior, where websites refuse to return content when HTTP requests include AI crawler user agents, and disclose that both categories of websites utilize it. Notably, the behavior of reputable news websites in this regard aligns more closely with their declared robots [.]txt directive than with that of misinformation websites. Finally, our longitudinal analysis reveals that this gap has widened over time, with AI-blocking by reputable sites increasing from 23% in September 2023 to nearly 60% by May 2025. Our findings highlight a growing asymmetry in content accessibility that may shape the training data available to LLMs, raising essential questions for web transparency, data ethics, and the future of AI training practices. CCS Concepts • Information systems → Web mining; Data exchange.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e085976f-424a-4cfb-956e-d75355d0bc62Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski et al.NDSS 2019 · 826 citations
- Who Funds Misinformation? A Systematic Analysis of the Ad-related Profit Routines of Fake News SitesEmmanouil Papadogiannakis, Panagiotis Papadopoulos, Evangelos P. Markatos, Nicolas KourtellisWWW 2023 · 39 citations
Related papers
- The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model ServicesJian Cui, Mingming Zha, XiaoFeng Wang, Xiaojing LiaoCCS 2025
- How does Misinformation Affect Large Language Model Behaviors and Preferences?Miao Peng, Nuo Chen, Jianheng Tang, Jia LiACL 2025 · 2 citations
- Characterizing the Implementation of Censorship Policies in Chinese LLM ServicesAnna Ablove, Shreyas Chandrashekaran, Xiao Qiang, Roya EnsafiNDSS 2026 · 2 citations
- Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated CitationsSunhao Dai, Zhanshuo Cao, Wenjie Wang, Liang Pang et al.EMNLP 2025
- Information Retrieval Induced Safety Degradation in AI AgentsCheng Yu, Benedikt Stroebl, Diyi Yang, Orestis PapakyriakopoulosNeurIPS 2025 · 5 citations
