The Smallest Extraction Problem
Valerio Cetorelli, Paolo Atzeni, Valter Crescenzi, Franco Milicchio
摘要
We introduce landmark grammars , a new family of context-free grammars aimed at describing the HTML source code of pages published by large and templated websites and therefore at effectively tackling Web data extraction problems. Indeed, they address the inherent ambiguity of HTML, one of the main challenges of Web data extraction, which, despite over twenty years of research, has been largely neglected by the approaches presented in literature.
We then formalize the Smallest Extraction Problem (SEP), an optimization problem for finding the grammar of a family that best describes a set of pages and contextually extract their data.
Finally, we present an unsupervised learning algorithm to induce a landmark grammar from a set of pages sharing a common HTML template, and we present an automatic Web data extraction system. The experiments on consolidated benchmarks show that the approach can substantially contribute to improve the state-of-the-art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsZhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu 等WWW 2023 · 被引用 7 次
- Web Record Extraction with InvariantsZhijia Chen, Weiyi Meng, Eduard C. DragutVLDB 2023 · 被引用 5 次
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung 等SIGMOD 2026
相关 Paper
- Landmarks and regions: a robust approach to data extractionSuresh Parthasarathy, Lincy Pattanaik, Anirudh Khatry, Arun Iyer 等PLDI 2022 · 被引用 2 次
- Activity Grammars for Temporal Action SegmentationDayoung Gong, Joonseok Lee, Deunsol Jung, Suha Kwak 等NeurIPS 2023 · 被引用 17 次
- VLGrammar: Grounded Grammar Induction of Vision and LanguageYining Hong, Qing Li, Song-Chun Zhu, Siyuan HuangICCV 2021 · 被引用 28 次
- Glean: Structured Extractions from Templatic DocumentsSandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa 等VLDB 2021 · 被引用 17 次
- FreeDOM: A Transferable Neural Architecture for Structured Information Extraction on Web DocumentsBill Yuchen Lin, Ying Sheng, Nguyen Vo, Sandeep TataKDD 2020 · 被引用 31 次
