Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages
Ritesh Sarkhel, Binxuan Huang, Colin Lockard, Prashant Shiralkar
摘要
Information Extraction (IE) from semi-structured web-pages is a long studied problem. Training a model for this extraction task requires a large number of human-labeled samples. Prior works have proposed transferable models to improve the label-efficiency of this training process. Extraction performance of transferable models however, depends on the size of their fine-tuning corpus. This holds true for large language models (LLM) such as GPT-3 as well. Generalist models like LLMs need to be fine-tuned on in-domain, human-labeled samples for competitive performance on this extraction task. Constructing a large-scale fine-tuning corpus with human-labeled samples, however, requires significant effort. In this paper, we develop a Label-Efficient Self-Training Algorithm (LEAST) to improve the label-efficiency of this fine-tuning process. Our contributions are two-fold. First , we develop a generative model that facilitates the construction of a large-scale fine-tuning corpus with minimal human-effort. Second , to ensure that the extraction performance does not suffer due to noisy training samples in our fine-tuning corpus, we develop an uncertainty-aware training strategy. Experiments on two publicly available datasets show that LEAST generalizes to multiple verticals and backbone models. Using LEAST, we can train models with less than ten human-labeled pages from each website, outperforming strong baselines while reducing the number of human-labeled training samples needed for comparable performance by up to 11 x.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim 等EMNLP 2022 · 被引用 285 次
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 被引用 182 次
- Learning to Extract Attribute Value from Product via Question Answering: A Multi-task ApproachQifan Wang, Li Yang, Bhargav Kanagal, Sumit Sanghai 等KDD 2020 · 被引用 75 次
相关 Paper
- STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language ModelsMingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham 等AAAI 2024 · 被引用 22 次
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LMMengjie Liu, Jiahui Peng, Wenchang Ning, Pei Chu 等KDD 2026 · 被引用 10 次
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 被引用 4 次
- Not All Documents Are What You Need for Extracting Instruction Tuning DataChi Zhang, Huaping Zhong, Hongtao Li, Chengliang Chai 等ICLR 2026 · 被引用 3 次
