Detecting and Diagnosing Errors in Serving Archived Web Pages
Jingyuan Zhu, Huanchen Sun, Harsha V. Madhyastha
摘要
Web archives crawl and save copies of pages from the web, enabling users to interact with web pages in the form they existed in the past. Prior to serving any archived page, an archive rewrites the page’s source so that users’ browsers fetch the page’s resources from the archive, not from the servers which originally hosted them. But, on many modern pages, an archive’s edits to crawled scripts result in a loss of fidelity, i.e., an archived copy fails to accurately mimic the original page even when the archive had crawled all resources on the page. To help the developers of archival systems identify and fix the bugs which result in incorrect rewrites of crawled pages, we present FidEx. First, FidEx enables accurate identification of the pages on which an archive violates fidelity. It does so by tracking and comparing the execution of scripts between when a page is crawled and when its copy is loaded. In comparison to existing methods which compare the two loads using either screenshots or the errors reported by the browser, FidEx reduces the false positive rate from around 70% to less than 10%. Second, on every page on which it identifies a loss of fidelity, FidEx pinpoints which subset of the archive’s edits to the page are erroneous. Leveraging this input to fix bugs in the most widely used archival system, we reduced the fraction of archived pages which violate fidelity from 15% to 9%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski 等NDSS 2019 · 被引用 826 次
- Fault Localization with Code Coverage Representation LearningYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 被引用 120 次
- Large Language Models for Test-Free Fault LocalizationAidan Z. H. Yang, Claire Le Goues, Ruben Martins, Vincent J. HellendoornICSE 2024 · 被引用 98 次
- R2Z2: Detecting Rendering Regressions in Web Browsers through Differential Fuzz TestingSuhwan Song, Jaewon Hur, Sunwoo Kim, Philip Rogers 等ICSE 2022 · 被引用 12 次
- Sprinter: Speeding Up High-Fidelity Crawling of the Modern WebAyush Goel, Jingyuan Zhu, Ravi Netravali, Harsha V. MadhyasthaNSDI 2024 · 被引用 5 次
相关 Paper
- Jawa: Web Archival in the Era of JavaScriptAyush Goel, Jingyuan Zhu, Ravi Netravali, Harsha V. MadhyasthaOSDI 2022
- Rewriting History: Changing the Archived Web from the PresentAda Lerner, Tadayoshi Kohno, Franziska RoesnerCCS 2017 · 被引用 35 次
- The Power to Never Be Wrong: Evasions and Anachronistic Attacks Against Web ArchivesRobin Kirchner, Chris Tsoukaladelis, Martin Johns, Nick NikiforakisCCS 2025
- FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented GenerationZhiping Zhou, Xiaohong Li, Ruitao Feng, Yao Zhang 等NDSS 2026 · 被引用 5 次
- Out of Sight, Out of Mind: Detecting Orphaned Web Pages at Internet-ScaleStijn Pletinckx, Kevin Borgolte, Tobias FiebigCCS 2021 · 被引用 11 次
