You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements
Xuenan Zhang, Yuqing Yang, Giancarlo Pellegrino
摘要
Web measurement studies rely on domain datasets such as Tranco to quantify the prevalence and impact of security issues at scale, but exhaustively analyzing these datasets is often infeasible because of the cost of advanced analysis techniques, requiring the use of sampling. Despite its widespread use, sampling remains largely guided by convention-most commonly Top 𝑁 domain selection-rather than evidence, and its influence on the validity and generalizability of security findings has received little systematic evaluation. Consequently, it remains unclear whether common sampling strategies introduce systematic bias, distort observed vulnerability rates, or limit comparability across studies.
In this work, we undertake, to the best of our knowledge, the first comprehensive investigation into how sampling methodologies affect the measurements and the conclusions. Through a comprehensive literature review and large-scale measurements of 500k Tranco and 24.8M Common Crawl hosts, we perform a comparative evaluation of datasets and sampling strategies. We show that, while Top 𝑁 sampling may be a rational strategy, the researchers have to bear in mind that Top 𝑁 does not reflect the overall distribution of the web. Instead, probability-based strategies yield stable, unbiased estimates for prevalence and many impact objectives. Hybrid sampling provides no advantages over pure probability sampling, as its deterministic prefix consistently contributes negatively to accuracy. Building on these results, we provide data-backed guidance for future studies, proposing to use an adaptive probability-based sampling strategy that remains effective even when the prevalence of the target issue is unknown. 1
• Security and privacy → Web application security; • General and reference → Measurement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski 等NDSS 2019 · 被引用 826 次
- How Well Do My Results Generalize? Comparing Security and Privacy Survey Results from MTurk, Web, and Telephone SamplesElissa M. Redmiles, Sean Kross, Michelle L. MazurekS&P 2019 · 被引用 222 次
- Don't Trust The Locals: Investigating the Prevalence of Persistent Client-Side Cross-Site Scripting in the WildMarius Steffens, Christian Rossow, Martin Johns, Ben StockNDSS 2019 · 被引用 84 次
- Riding out DOMsday: Towards Detecting and Preventing DOM Cross-Site ScriptingWilliam Melicher, Anupam Das, Mahmood Sharif, Lujo Bauer 等NDSS 2018 · 被引用 84 次
- Content Security Problems?: Evaluating the Effectiveness of Content Security Policy in the WildStefano Calzavara, Alvise Rabitti, Michele BugliesiCCS 2016 · 被引用 71 次
相关 Paper
- Towards Realistic and ReproducibleWeb Crawl MeasurementsJordan Jueckstock, Shaown Sarker, Peter Snyder, Aidan Beggs 等WWW 2021 · 被引用 52 次
- Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the WebSyed Suleman Ahmad, Muhammad Daniyal Dar, Muhammad Fareed Zaffar, Narseo Vallina-Rodriguez 等WWW 2020 · 被引用 42 次
- SoK: State of the Krawlers - Evaluating the Effectiveness of Crawling Algorithms for Web Security MeasurementsAleksei Stafeev, Giancarlo PellegrinoUSENIX Security 2024 · 被引用 12 次
- Can I Take Your Subdomain? Exploring Same-Site Attacks in the Modern WebMarco Squarcina, Mauro Tempesta, Lorenzo Veronese, Stefano Calzavara 等USENIX Security 2021 · 被引用 30 次
- Out of Sight, Out of Mind: Detecting Orphaned Web Pages at Internet-ScaleStijn Pletinckx, Kevin Borgolte, Tobias FiebigCCS 2021 · 被引用 11 次
