Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource Settings
Zhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu, Yubin Wang, Li Guo
摘要
Extracting structured information from all manner of webpages is an important problem with the potential to automate many real-world applications. Recent work has shown the effectiveness of leveraging DOM trees and pre-trained language models to describe and encode webpages. However, they typically optimize the model to learn the semantic co-occurrence of elements and labels in the same webpage, thus their effectiveness depends on sufficient labeled data, which is labor-intensive. In this paper, we further observe structural co-occurrences in different webpages of the same website: the same position in the DOM tree usually plays the same semantic role, and the DOM nodes in this position also share similar surface forms. Motivated by this, we propose a novel method, Structor, to effectively incorporate the structural co-occurrences over DOM tree and surface form into pre-trained language models. Such structural co-occurrences help the model learn the task better under low-resource settings, and we study two challenging experimental scenarios: website-level low-resource setting and webpage-level low-resource setting, to evaluate our approach. Extensive experiments on the public SWDE dataset show that Structor significantly outperforms the state-of-the-art models in both settings, and even achieves three times the performance of the strong baseline model in the case of extreme lack of training data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang 等AAAI 2020 · 被引用 898 次
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language ModelWenhan Xiong, Jingfei Du, William Yang Wang, Veselin StoyanovICLR 2020 · 被引用 215 次
- Reinforced Negative Sampling over Knowledge Graph for RecommendationXiang Wang, Yaokun Xu, Xiangnan He, Yixin Cao 等WWW 2020 · 被引用 209 次
- Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NERDong-Ho Lee, Akshen Kadakia, Kangmin Tan, Mahak Agarwal 等ACL 2022 · 被引用 96 次
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng 等WWW 2022 · 被引用 88 次
相关 Paper
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian 等SIGIR 2022 · 被引用 30 次
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng 等ACL 2023 · 被引用 16 次
- WIERT: Web Information Extraction via Render TreeZimeng Li, Bo Shao, Linjun Shou, Ming Gong 等AAAI 2023 · 被引用 10 次
- StructuralLM: Structural Pre-training for Form UnderstandingChenliang Li, Bin Bi, Ming Yan, Wei Wang 等ACL 2021
- STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language ModelsMingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham 等AAAI 2024 · 被引用 22 次
