Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the Web
Syed Suleman Ahmad, Muhammad Daniyal Dar, Muhammad Fareed Zaffar, Narseo Vallina-Rodriguez, Rishab Nithyanand
摘要
Data generated by web crawlers has formed the basis for much of our current understanding of the Internet. However, not all crawlers are created equal and crawlers generally find themselves trading off between computational overhead, developer effort, data accuracy, and completeness. Therefore, the choice of crawler has a critical impact on the data generated and knowledge inferred from it. In this paper, we conduct a systematic study of the trade-offs presented by different crawlers and the impact that these can have on various types of measurement studies. We make the following contributions: First, we conduct a survey of all research published since 2015 in the premier security and Internet measurement venues to identify and verify the repeatability of crawling methodologies deployed for different problem domains and publication venues. Next, we conduct a qualitative evaluation of a subset of all crawling tools identified in our survey. This evaluation allows us to draw conclusions about the suitability of each tool for specific types of data gathering. Finally, we present a methodology and a measurement framework to empirically highlight the differences between crawlers and how the choice of crawler can impact our understanding of the web.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Towards Realistic and ReproducibleWeb Crawl MeasurementsJordan Jueckstock, Shaown Sarker, Peter Snyder, Aidan Beggs 等WWW 2021 · 被引用 52 次
- Reproducibility and Replicability of Web Measurement StudiesNurullah Demir, Matteo Große-Kampmann, Tobias Urban, Christian Wressnegger 等WWW 2022 · 被引用 47 次
- SoK: State of the Krawlers - Evaluating the Effectiveness of Crawling Algorithms for Web Security MeasurementsAleksei Stafeev, Giancarlo PellegrinoUSENIX Security 2024 · 被引用 12 次
- To Auth or Not To Auth? A Comparative Analysis of the Pre- and Post-Login Security LandscapeJannis Rautenstrauch, Metodi Mitkov, Thomas Helbrecht, Lorenz Hetterich 等S&P 2024 · 被引用 6 次
- You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security MeasurementsFlorian Hantke, Stefano Calzavara, Moritz Wilhelm, Alvise Rabitti 等CCS 2023 · 被引用 4 次
它引用的顶会 Paper7
- Online Tracking: A 1-million-site Measurement and AnalysisSteven Englehardt, Arvind NarayananCCS 2016 · 被引用 798 次
- Internet Jones and the Raiders of the Lost Trackers: An Archaeological Study of Web Tracking from 1996 to 2016Ada Lerner, Anna Kornfeld Simpson, Tadayoshi Kohno, Franziska RoesnerUSENIX Security 2016 · 被引用 273 次
- The use of TLS in Censorship CircumventionSergey Frolov, Eric WustrowNDSS 2019 · 被引用 97 次
- ICLab: A Global, Longitudinal Internet Censorship Measurement PlatformArian Akhavan Niaki, Shinyoung Cho, Zachary Weinberg, Nguyen Phong Hoang 等S&P 2020 · 被引用 94 次
- The Web's Sixth Sense: A Study of Scripts Accessing Smartphone SensorsAnupam Das, Gunes Acar, Nikita Borisov, Amogh PradeepCCS 2018 · 被引用 91 次
相关 Paper
- Web Execution Bundles: Reproducible, Accurate, and Archivable Web MeasurementsFlorian Hantke, Peter Snyder, Hamed Haddadi, Ben StockUSENIX Security 2025
- You Get What You Sample: Evaluating Sampling Strategies for Web Security MeasurementsXuenan Zhang, Yuqing Yang, Giancarlo PellegrinoCCS 2026
- The Representativeness of Automated Web Crawls as a Surrogate for Human BrowsingDavid Zeber, Sarah Bird, Camila Oliveira, Walter Rudametkin 等WWW 2020 · 被引用 37 次
- SoK: After Decades of Web Tracker Detection, What's Next?Wolf Rieder, Philip Raschke, Thomas Cory, Christian René Sechting 等S&P 2026 · 被引用 1 次
- Sprinter: Speeding Up High-Fidelity Crawling of the Modern WebAyush Goel, Jingyuan Zhu, Ravi Netravali, Harsha V. MadhyasthaNSDI 2024 · 被引用 5 次
