Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource Settings
Zhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu, Yubin Wang, Li Guo
Abstract
Extracting structured information from all manner of webpages is an important problem with the potential to automate many real-world applications. Recent work has shown the effectiveness of leveraging DOM trees and pre-trained language models to describe and encode webpages. However, they typically optimize the model to learn the semantic co-occurrence of elements and labels in the same webpage, thus their effectiveness depends on sufficient labeled data, which is labor-intensive. In this paper, we further observe structural co-occurrences in different webpages of the same website: the same position in the DOM tree usually plays the same semantic role, and the DOM nodes in this position also share similar surface forms. Motivated by this, we propose a novel method, Structor, to effectively incorporate the structural co-occurrences over DOM tree and surface form into pre-trained language models. Such structural co-occurrences help the model learn the task better under low-resource settings, and we study two challenging experimental scenarios: website-level low-resource setting and webpage-level low-resource setting, to evaluate our approach. Extensive experiments on the public SWDE dataset show that Structor significantly outperforms the state-of-the-art models in both settings, and even achieves three times the performance of the strong baseline model in the case of extreme lack of training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c2fa2b1-ea63-4e08-bf41-b232ec6fe4d6Builds on13
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang et al.AAAI 2020 · 898 citations
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language ModelWenhan Xiong, Jingfei Du, William Yang Wang, Veselin StoyanovICLR 2020 · 215 citations
- Reinforced Negative Sampling over Knowledge Graph for RecommendationXiang Wang, Yaokun Xu, Xiangnan He, Yixin Cao et al.WWW 2020 · 209 citations
- Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NERDong-Ho Lee, Akshen Kadakia, Kangmin Tan, Mahak Agarwal et al.ACL 2022 · 96 citations
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng et al.WWW 2022 · 88 citations
Related papers
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian et al.SIGIR 2022 · 30 citations
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng et al.ACL 2023 · 16 citations
- WIERT: Web Information Extraction via Render TreeZimeng Li, Bo Shao, Linjun Shou, Ming Gong et al.AAAI 2023 · 10 citations
- StructuralLM: Structural Pre-training for Form UnderstandingChenliang Li, Bin Bi, Ming Yan, Wei Wang et al.ACL 2021
- STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language ModelsMingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham et al.AAAI 2024 · 22 citations
