Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages
Ritesh Sarkhel, Binxuan Huang, Colin Lockard, Prashant Shiralkar
Abstract
Information Extraction (IE) from semi-structured web-pages is a long studied problem. Training a model for this extraction task requires a large number of human-labeled samples. Prior works have proposed transferable models to improve the label-efficiency of this training process. Extraction performance of transferable models however, depends on the size of their fine-tuning corpus. This holds true for large language models (LLM) such as GPT-3 as well. Generalist models like LLMs need to be fine-tuned on in-domain, human-labeled samples for competitive performance on this extraction task. Constructing a large-scale fine-tuning corpus with human-labeled samples, however, requires significant effort. In this paper, we develop a Label-Efficient Self-Training Algorithm (LEAST) to improve the label-efficiency of this fine-tuning process. Our contributions are two-fold. First , we develop a generative model that facilitates the construction of a large-scale fine-tuning corpus with minimal human-effort. Second , to ensure that the extraction performance does not suffer due to noisy training samples in our fine-tuning corpus, we develop an uncertainty-aware training strategy. Experiments on two publicly available datasets show that LEAST generalizes to multiple verticals and backbone models. Using LEAST, we can train models with less than ten human-labeled pages from each website, outperforming strong baselines while reducing the number of human-labeled training samples needed for comparable performance by up to 11 x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 718bc7b2-045f-4252-aa84-3558ebe94732Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim et al.EMNLP 2022 · 285 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
- Learning to Extract Attribute Value from Product via Question Answering: A Multi-task ApproachQifan Wang, Li Yang, Bhargav Kanagal, Sumit Sanghai et al.KDD 2020 · 75 citations
Related papers
- STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language ModelsMingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham et al.AAAI 2024 · 22 citations
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LMMengjie Liu, Jiahui Peng, Wenchang Ning, Pei Chu et al.KDD 2026 · 10 citations
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose et al.EMNLP 2021 · 97 citations
- Entity Extraction in Low Resource Domains with Selective Pre-training of Large Language ModelsAniruddha Mahapatra, Sharmila Reddy Nangi, Aparna Garimella, Anandhavelu NatarajanEMNLP 2022 · 4 citations
- Not All Documents Are What You Need for Extracting Instruction Tuning DataChi Zhang, Huaping Zhong, Hongtao Li, Chengliang Chai et al.ICLR 2026 · 3 citations
