BAND: Biomedical Alert News Dataset
Zihao Fu, Meiru Zhang, Zaiqiao Meng, Yannan Shen, David L. Buckeridge, Nigel Collier
摘要
Infectious disease outbreaks continue to pose a significant threat to human health and well-being. To improve disease surveillance and understanding of disease spread, several surveillance systems have been developed to monitor daily news alerts and social media. However, existing systems lack thorough epidemiological analysis in relation to corresponding alerts or news, largely due to the scarcity of well-annotated reports data. To address this gap, we introduce the Biomedical Alert News Dataset (BAND), which includes 1,508 samples from existing reported news articles, open emails, and alerts, as well as 30 epidemiology-related questions. These questions necessitate the model's expert reasoning abilities, thereby offering valuable insights into the outbreak of the disease. The BAND dataset brings new challenges to the NLP world, requiring better inference capability of the content and the ability to infer important information. We provide several benchmark tasks, including Named Entity Recognition (NER), Question Answering (QA), and Event Extraction (EE), to demonstrate existing models' capabilities and limitations in handling epidemiology-specific tasks. It is worth noting that some models may lack the human-like inference capability required to fully utilize the corpus. To the best of our knowledge, the BAND corpus is the largest corpus of well-annotated biomedical outbreak alert news with elaborately designed questions, making it a valuable resource for epidemiologists and NLP researchers alike.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Prefix-Tuning: Optimizing Continuous Prompts for GenerationXiang Lisa Li, Percy LiangACL 2021
相关 Paper
- ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital EnvironmentsSourjyadip Ray, Kushal Gupta, Soumi Kundu, Payal Arvind Kasat 等EMNLP 2024 · 被引用 1 次
- COMETA: A Corpus for Medical Entity Linking in the Social MediaMarco Basaldella, Fangyu Liu, Ehsan Shareghi, Nigel CollierEMNLP 2020 · 被引用 4 次
- SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLPDecheng Duan, Jitong Peng, Yingyi Zhang, Chengzhi ZhangEMNLP 2025 · 被引用 1 次
- PHEE: A Dataset for Pharmacovigilance Event Extraction from TextZhaoyue Sun, Jiazheng Li, Gabriele Pergola, Byron C. Wallace 等EMNLP 2022 · 被引用 17 次
- MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive AnnotationsVishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu 等EMNLP 2024 · 被引用 1 次
