Sprinter: Speeding Up High-Fidelity Crawling of the Modern Web
Ayush Goel, Jingyuan Zhu, Ravi Netravali, Harsha V. Madhyastha
摘要
Crawling the web at scale forms the basis of many important systems: web search engines, smart assistants, generative AI, web archives, and so on. Yet, the research community has paid little attention to this workload in the last decade. In this paper, we highlight the need to revisit the notion that web crawling is a solved problem. Specifically, to discover and fetch all page resources dependent on JavaScript and modern web APIs, crawlers today have to employ compute-intensive web browsers. This significantly inflates the scale of the infrastructure necessary to crawl pages at high throughput.
To make web crawling more efficient without any loss of fidelity, we present Sprinter, which combines browser-based and browserless crawling to get the best of both. The key to Sprinter's design is our observation that crawling workloads typically include many pages per site and, unlike in traditional user-facing page loads, there is significant potential to reuse client-side computations across pages. Taking advantage of this property, Sprinter crawls a small, carefully chosen, subset of pages on each site using a browser, and then efficiently identifies and exploits opportunities to reuse the browser's computations on other pages. Sprinter was able to crawl a corpus of 50,000 pages 5x faster than browser-based crawling, while still closely matching a browser in the set of resources fetched.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin 等OSDI 2022 · 被引用 105 次
- CSPAutoGen: Black-box Enforcement of Content Security Policy upon Real-world WebsitesXiang Pan, Yinzhi Cao, Shuangping Liu, Yu Zhou 等CCS 2016 · 被引用 55 次
- Oblique: Accelerating Page Loads Using Symbolic ExecutionRonny Ko, James Mickens, Blake Loring, Ravi NetravaliNSDI 2021 · 被引用 13 次
相关 Paper
- Theseus: Smart Web Crawling via Resource-Guided Semantic ModelingYongheng Huang, Chenghang Shi, Wenxiao Yao, Jie Lu 等CCS 2026
- The Representativeness of Automated Web Crawls as a Surrogate for Human BrowsingDavid Zeber, Sarah Bird, Camila Oliveira, Walter Rudametkin 等WWW 2020 · 被引用 37 次
- Reproducibility and Replicability of Web Measurement StudiesNurullah Demir, Matteo Große-Kampmann, Tobias Urban, Christian Wressnegger 等WWW 2022 · 被引用 47 次
- Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the WebSyed Suleman Ahmad, Muhammad Daniyal Dar, Muhammad Fareed Zaffar, Narseo Vallina-Rodriguez 等WWW 2020 · 被引用 42 次
- Horcrux: Automatic JavaScript Parallelism for Resource-Efficient Web ComputationShaghayegh Mardani, Ayush Goel, Ronny Ko, Harsha V. Madhyastha 等OSDI 2021 · 被引用 9 次
