Web Record Extraction with Invariants
Zhijia Chen, Weiyi Meng, Eduard C. Dragut
Abstract
Web records are structured data on a Web page that embeds records retrieved from an underlying database according to some templates. Mining data records on the Web enables the integration of data from multiple Web sites for providing value-added services. Most existing works on Web record extraction make two key assumptions: (1) records are retrieved from databases with uniform schemas and (2) records are displayed in a linear structure on a Web page. These assumptions no longer hold on the modern Web. A Web page may present records of diverse entity types with different schemas and organize records hierarchically, in nested structures, to show richer relationships among records. In this paper, we revisit these assumptions and modify them to reflect Web pages on the modern Web. Based on the reformulated assumptions, we introduce the concept of invariant in Web data records and propose Miria ( Mi ning r ecord i nvari a nt), a bottom-up, recursive approach to construct the Web records from the invariants. The proposed approach is both effective and efficient, consistently outperforming the state-of-the-art Web record extraction methods on modern Web pages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea et al.EMNLP 2024 · 7 citations
- StructVizor: Interactive Profiling of Semi-Structured Textual DataYanwei Huang, Yan Miao, Di Weng, Adam Perer et al.CHI 2025 · 3 citations
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung et al.SIGMOD 2026
Builds on1
Related papers
- MINES: Explainable Anomaly Detection through Web API Invariant InferenceWenjie Zhang, Yun Lin, Chun Fung Amos Kwok, Xiwen Teoh et al.ICSE 2026
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement LearningShicheng Liu, Kai Sun, Lisheng Fu, Xilun Chen et al.ICLR 2026 · 2 citations
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng et al.ACL 2023 · 16 citations
- Hierarchical Entity Resolution using an OracleSainyam Galhotra, Donatella Firmani, Barna Saha, Divesh SrivastavaSIGMOD 2022 · 4 citations
- Dual-View Visual Contextualization for Web NavigationJihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng et al.CVPR 2024
