Generating Diverse Training Samples for Relation Extraction with Large Language Models
Zexuan Li, Hongliang Dai, Piji Li
摘要
Using Large Language Models (LLMs) to generate training data can potentially be a preferable way to improve zero or few-shot NLP tasks. However, many problems remain to be investigated for this direction. For the task of Relation Extraction (RE), we find that samples generated by directly prompting LLMs may easily have high structural similarities with each other. They tend to use a limited variety of phrasing while expressing the relation between a pair of entities. Therefore, in this paper, we study how to effectively improve the diversity of the training samples generated with LLMs for RE, while also maintaining their correctness. We first try to make the LLMs produce dissimilar samples by directly giving instructions in In-Context Learning (ICL) prompts. Then, we propose an approach to fine-tune LLMs for diversity training sample generation through Direct Preference Optimization (DPO). Our experiments on commonly used RE datasets show that both attempts can improve the quality of the generated training data. We also find that comparing with directly performing RE with an LLM, training a non-LLM RE model with its generated samples may lead to better performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation ExtractionAunabil Chakma, Mihai Surdeanu, Eduardo BlancoACL 2026 · 被引用 1 次
- Learning to Generate and Extract: A Multi-Agent Collaboration Framework for Zero-Shot Document-Level Event Arguments ExtractionGuangjun Zhang, Hu Zhang, Yazhou Han, Yue Fan 等AAAI 2026
- M-BRe: Discovering Training Samples for Relation Extraction from Unlabeled Texts with Large Language ModelsZexuan Li, Hongliang Dai, Piji LiEMNLP 2025
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation ExtractionXiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng 等WWW 2022 · 被引用 488 次
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 被引用 309 次
相关 Paper
- CodeIE: Large Code Generation Models are Better Few-Shot Information ExtractorsPeng Li, Tianxiang Sun, Qiong Tang, Hang Yan 等ACL 2023 · 被引用 41 次
- Grasping the Essentials: Tailoring Large Language Models for Zero-Shot Relation ExtractionSizhe Zhou, Yu Meng, Bowen Jin, Jiawei HanEMNLP 2024 · 被引用 6 次
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- Topic-Oriented Open Relation Extraction with A Priori Seed GenerationLinyi Ding, Jinfeng Xiao, Sizhe Zhou, Chaoqi Yang 等EMNLP 2024 · 被引用 2 次
- On the Role of Discriminative Models in Generative Relation ExtractionGuozheng Li, Peng Wang, Zijie Xu, Jing Zhou 等ACL 2026
