Detecting and Diagnosing Errors in Serving Archived Web Pages
Jingyuan Zhu, Huanchen Sun, Harsha V. Madhyastha
Abstract
Web archives crawl and save copies of pages from the web, enabling users to interact with web pages in the form they existed in the past. Prior to serving any archived page, an archive rewrites the page’s source so that users’ browsers fetch the page’s resources from the archive, not from the servers which originally hosted them. But, on many modern pages, an archive’s edits to crawled scripts result in a loss of fidelity, i.e., an archived copy fails to accurately mimic the original page even when the archive had crawled all resources on the page. To help the developers of archival systems identify and fix the bugs which result in incorrect rewrites of crawled pages, we present FidEx. First, FidEx enables accurate identification of the pages on which an archive violates fidelity. It does so by tracking and comparing the execution of scripts between when a page is crawled and when its copy is loaded. In comparison to existing methods which compare the two loads using either screenshots or the errors reported by the browser, FidEx reduces the false positive rate from around 70% to less than 10%. Second, on every page on which it identifies a loss of fidelity, FidEx pinpoints which subset of the archive’s edits to the page are erroneous. Leveraging this input to fix bugs in the most widely used archival system, we reduced the fraction of archived pages which violate fidelity from 15% to 9%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ff986c0-bdf7-4fed-8f09-6dfeff503565Builds on6
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski et al.NDSS 2019 · 826 citations
- Fault Localization with Code Coverage Representation LearningYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 120 citations
- Large Language Models for Test-Free Fault LocalizationAidan Z. H. Yang, Claire Le Goues, Ruben Martins, Vincent J. HellendoornICSE 2024 · 98 citations
- R2Z2: Detecting Rendering Regressions in Web Browsers through Differential Fuzz TestingSuhwan Song, Jaewon Hur, Sunwoo Kim, Philip Rogers et al.ICSE 2022 · 12 citations
- Sprinter: Speeding Up High-Fidelity Crawling of the Modern WebAyush Goel, Jingyuan Zhu, Ravi Netravali, Harsha V. MadhyasthaNSDI 2024 · 5 citations
Related papers
- Jawa: Web Archival in the Era of JavaScriptAyush Goel, Jingyuan Zhu, Ravi Netravali, Harsha V. MadhyasthaOSDI 2022
- Rewriting History: Changing the Archived Web from the PresentAda Lerner, Tadayoshi Kohno, Franziska RoesnerCCS 2017 · 35 citations
- The Power to Never Be Wrong: Evasions and Anachronistic Attacks Against Web ArchivesRobin Kirchner, Chris Tsoukaladelis, Martin Johns, Nick NikiforakisCCS 2025
- FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented GenerationZhiping Zhou, Xiaohong Li, Ruitao Feng, Yao Zhang et al.NDSS 2026 · 5 citations
- Out of Sight, Out of Mind: Detecting Orphaned Web Pages at Internet-ScaleStijn Pletinckx, Kevin Borgolte, Tobias FiebigCCS 2021 · 11 citations
