S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation Extraction
Benfeng Xu, Quan Wang, Yajuan Lyu, Dai Dai, Yongdong Zhang, Zhendong Mao
Abstract
Current relation extraction methods suffer from the inadequacy of large-scale annotated data. While distant supervision alleviates the problem of data quantities, there still exists domain disparity in data qualities due to its reliance on domain-restrained knowledge bases. In this work, we propose S 2 ynRE, a framework of two-stage Self-training with Synthetic data for Relation Extraction. We first leverage the capability of large language models to adapt to the target domain and automatically synthesize large quantities of coherent, realistic training data. We then propose an accompanied two-stage self-training algorithm that iteratively and alternately learns from synthetic and golden data together. We conduct comprehensive experiments and detailed ablations on popular relation extraction datasets to demonstrate the effectiveness of the proposed framework. Code is available at https: //github.com/BenfengXu/S2ynRE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7965cf1-895e-48f3-9c26-ff878235728eCited by top-tier papers5
- Reward-based Input Construction for Cross-document Relation ExtractionByeonghu Na, Suhyeon Jo, Yeongmin Kim, Il-Chul MoonACL 2024 · 3 citations
- Data-Constrained Synthesis of Training Data for De-IdentificationThomas Vakili, Aron Henriksson, Hercules DalianisACL 2025 · 3 citations
- Understanding Synthetic Context Extension via Retrieval HeadsXinyu Zhao, Fangcong Yin, Greg DurrettICML 2025
- SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on SchemaYushen Fang, Jianjun Li, Mingqian Ding, Chang Liu et al.AAAI 2026
- When Phrases Meet Probabilities: Enabling Open Relation Extraction with Cooperating Large Language ModelsJiaxin Wang, Lingling Zhang, Wee Sun Lee, Yujie Zhong et al.ACL 2024
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation ExtractionXiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng et al.WWW 2022 · 488 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 309 citations
- Learning from Context or Names? An Empirical Study on Neural Relation ExtractionHao Peng, Tianyu Gao, Xu Han, Yankai Lin et al.EMNLP 2020 · 185 citations
Related papers
- Knowing False Negatives: An Adversarial Training Method for Distantly Supervised Relation ExtractionKailong Hao, Botao Yu, Wei HuEMNLP 2021 · 19 citations
- Revisiting the Negative Data of Distantly Supervised Relation ExtractionChenhao Xie, Jiaqing Liang, Jingping Liu, Chengsong Huang et al.ACL 2021
- SelfORE: Self-supervised Relational Feature Learning for Open Relation ExtractionXuming Hu, Lijie Wen, Yusong Xu, Chenwei Zhang et al.EMNLP 2020 · 81 citations
- Improving Distantly Supervised Relation Extraction by Natural Language InferenceKang Zhou, Qiao Qiao, Yuepei Li, Qi LiAAAI 2023 · 12 citations
- fmLRE: A Low-Resource Relation Extraction Model Based on Feature Mapping Similarity CalculationPeng Wang, Tong Shao, Ke Ji, Guozheng Li et al.AAAI 2023 · 8 citations
