The Smallest Extraction Problem
Valerio Cetorelli, Paolo Atzeni, Valter Crescenzi, Franco Milicchio
Abstract
We introduce landmark grammars , a new family of context-free grammars aimed at describing the HTML source code of pages published by large and templated websites and therefore at effectively tackling Web data extraction problems. Indeed, they address the inherent ambiguity of HTML, one of the main challenges of Web data extraction, which, despite over twenty years of research, has been largely neglected by the approaches presented in literature.
We then formalize the Smallest Extraction Problem (SEP), an optimization problem for finding the grammar of a family that best describes a set of pages and contextually extract their data.
Finally, we present an unsupervised learning algorithm to induce a landmark grammar from a set of pages sharing a common HTML template, and we present an automatic Web data extraction system. The experiments on consolidated benchmarks show that the approach can substantially contribute to improve the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsZhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu et al.WWW 2023 · 7 citations
- Web Record Extraction with InvariantsZhijia Chen, Weiyi Meng, Eduard C. DragutVLDB 2023 · 5 citations
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung et al.SIGMOD 2026
Related papers
- Landmarks and regions: a robust approach to data extractionSuresh Parthasarathy, Lincy Pattanaik, Anirudh Khatry, Arun Iyer et al.PLDI 2022 · 2 citations
- Activity Grammars for Temporal Action SegmentationDayoung Gong, Joonseok Lee, Deunsol Jung, Suha Kwak et al.NeurIPS 2023 · 17 citations
- VLGrammar: Grounded Grammar Induction of Vision and LanguageYining Hong, Qing Li, Song-Chun Zhu, Siyuan HuangICCV 2021 · 28 citations
- Glean: Structured Extractions from Templatic DocumentsSandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa et al.VLDB 2021 · 17 citations
- FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web DocumentsBill Yuchen Lin, Ying Sheng, Nguyen Vo, Sandeep TataKDD 2020 · 31 citations
