Generating Diverse Training Samples for Relation Extraction with Large Language Models
Zexuan Li, Hongliang Dai, Piji Li
Abstract
Using Large Language Models (LLMs) to generate training data can potentially be a preferable way to improve zero or few-shot NLP tasks. However, many problems remain to be investigated for this direction. For the task of Relation Extraction (RE), we find that samples generated by directly prompting LLMs may easily have high structural similarities with each other. They tend to use a limited variety of phrasing while expressing the relation between a pair of entities. Therefore, in this paper, we study how to effectively improve the diversity of the training samples generated with LLMs for RE, while also maintaining their correctness. We first try to make the LLMs produce dissimilar samples by directly giving instructions in In-Context Learning (ICL) prompts. Then, we propose an approach to fine-tune LLMs for diversity training sample generation through Direct Preference Optimization (DPO). Our experiments on commonly used RE datasets show that both attempts can improve the quality of the generated training data. We also find that comparing with directly performing RE with an LLM, training a non-LLM RE model with its generated samples may lead to better performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 072f8a77-29dc-4f3a-a772-c9ce89eb2cddCited by top-tier papers3
- Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation ExtractionAunabil Chakma, Mihai Surdeanu, Eduardo BlancoACL 2026 · 1 citation
- Learning to Generate and Extract: A Multi-Agent Collaboration Framework for Zero-Shot Document-Level Event Arguments ExtractionGuangjun Zhang, Hu Zhang, Yazhou Han, Yue Fan et al.AAAI 2026
- M-BRe: Discovering Training Samples for Relation Extraction from Unlabeled Texts with Large Language ModelsZexuan Li, Hongliang Dai, Piji LiEMNLP 2025
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation ExtractionXiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng et al.WWW 2022 · 488 citations
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 309 citations
Related papers
- CodeIE: Large Code Generation Models are Better Few-Shot Information ExtractorsPeng Li, Tianxiang Sun, Qiong Tang, Hang Yan et al.ACL 2023 · 41 citations
- Grasping the Essentials: Tailoring Large Language Models for Zero-Shot Relation ExtractionSizhe Zhou, Yu Meng, Bowen Jin, Jiawei HanEMNLP 2024 · 6 citations
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- Topic-Oriented Open Relation Extraction with A Priori Seed GenerationLinyi Ding, Jinfeng Xiao, Sizhe Zhou, Chaoqi Yang et al.EMNLP 2024 · 2 citations
- On the Role of Discriminative Models in Generative Relation ExtractionGuozheng Li, Peng Wang, Zijie Xu, Jing Zhou et al.ACL 2026
