STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language Models
Mingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham, Nanyun Peng, Wei Wang
Abstract
Information extraction tasks such as event extraction require an in-depth understanding of the output structure and sub-task dependencies. They heavily rely on task-specific training data in the form of (passage, target structure) pairs to obtain reasonable performance. However, obtaining such data through human annotation is costly, leading to a pressing need for low-resource information extraction approaches that require minimal human labeling for real-world applications. Fine-tuning supervised models with synthesized training data would be a generalizable method, but the existing data generation methods either still rely on large-scale ground-truth data or cannot be applied to complicated IE tasks due to their poor performance. To address these challenges, we propose STAR, a data generation method that leverages Large Language Models (LLMs) to synthesize data instances given limited seed demonstrations, thereby boosting low-resource information extraction performance. Our approach involves generating target structures (Y) followed by generating passages (X), all accomplished with the aid of LLMs. We design fine-grained step-by-step instructions to obtain the initial data instances. We further reduce errors and improve data quality through self-reflection error identification and self-refinement with iterative revision. Our experiments show that the data generated by STAR significantly improve the performance of low-resource event extraction and relation extraction tasks, even surpassing the effectiveness of human-curated data. Human assessment of the data quality shows STAR-generated data exhibit higher passage quality and better align with the task definitions compared with the human-curated data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3f19385-495d-4590-a969-659b02e80bd3Cited by top-tier papers5
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen et al.WWW 2025 · 16 citations
- Memorize and Rank: Elevating Large Language Models for Clinical Diagnosis PredictionMingyu Derek Ma, Xiaoxuan Wang, Yijia Xiao, Anthony Cuturrufo et al.AAAI 2025 · 7 citations
- SNaRe: Domain-aware Data Generation for Low-Resource Event DetectionTanmay Parekh, Yuxuan Dong, Lucas Bandarkar, Artin Kim et al.EMNLP 2025 · 1 citation
- DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM ReasoningTanmay Parekh, Kartik Mehta, Ninareh Mehrabi, Kai-Wei Chang et al.EMNLP 2025 · 1 citation
Builds on14
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- A Joint Neural Model for Information Extraction with Global FeaturesYing Lin, Heng Ji, Fei Huang, Lingfei WuACL 2020 · 376 citations
Related papers
- Is a Large Language Model a Good Annotator for Event Extraction?Ruirui Chen, Chengwei Qin, Weifeng Jiang, Dongkyu ChoiAAAI 2024 · 65 citations
- Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information ExtractionMartin Josifoski, Marija Sakota, Maxime Peyrard, Robert WestEMNLP 2023 · 43 citations
- Reliable Data Generation and Selection for Low-Resource Relation ExtractionJunjie Yu, Xing Wang, Wenliang ChenAAAI 2024 · 7 citations
- A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource LanguagesTatiana Anikina, Ján Cegin, Jakub Simko, Simon OstermannEMNLP 2025
- Bridging the Gap: Aligning Language Model Generation with Structured Information Extraction via Controllable State TransitionHao Li, Yubing Ren, Yanan Cao, Yingjie Li et al.WWW 2025
