SNaRe: Domain-aware Data Generation for Low-Resource Event Detection
Tanmay Parekh, Yuxuan Dong, Lucas Bandarkar, Artin Kim, I-Hung Hsu, Kai-Wei Chang, Nanyun Peng
摘要
Event Detection (ED) -the task of identifying event mentions from natural language text -is critical for enabling reasoning in highly specialized domains such as biomedicine, law, and epidemiology. Data generation has proven to be effective in broadening its utility to wider applications without requiring expensive expert annotations. However, when existing generation approaches are applied to specialized domains, they struggle with label noise, where annotations are incorrect, and domain drift, characterized by a distributional mismatch between generated sentences and the target domain. To address these issues, we introduce SNARE, a domain-aware synthetic data generation framework composed of three components: Scout, Narrator, and Refiner. Scout extracts triggers from unlabeled target domain data and curates a high-quality domain-specific trigger list using corpus-level statistics to mitigate domain drift. Narrator, conditioned on these triggers, generates high-quality domainaligned sentences, and Refiner identifies additional event mentions, ensuring high annotation quality. Experimentation on three diverse domain ED datasets reveals how SNARE outperforms the best baseline, achieving average F1 gains of 3-7% in the zero-shot/few-shot settings and 4-20% F1 improvement for multilingual generation. Analyzing the generated trigger hit rate and human evaluation substantiates SNARE's stronger annotation quality and reduced domain drift. We will release our code at https://github.com/PlusLabNLP/SNaRe .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
相关 Paper
- DICE: Data-Efficient Clinical Event Extraction with Generative ModelsMingyu Derek Ma, Alexander Taylor, Wei Wang, Nanyun PengACL 2023 · 被引用 18 次
- OntoED: Low-resource Event Detection with Ontology EmbeddingShumin Deng, Ningyu Zhang, Luoqiu Li, Hui Chen 等ACL 2021
- Improving Event Detection via Open-domain Trigger KnowledgeMeihan Tong, Bin Xu, Shuai Wang, Yixin Cao 等ACL 2020 · 被引用 107 次
- BioFEG: Generate Latent Features for Biomedical Entity LinkingXuhui Sui, Ying Zhang, Xiangrui Cai, Kehui Song 等EMNLP 2023 · 被引用 6 次
- MAVEN: A Massive General Domain Event Detection DatasetXiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang 等EMNLP 2020 · 被引用 143 次
