BAND: Biomedical Alert News Dataset
Zihao Fu, Meiru Zhang, Zaiqiao Meng, Yannan Shen, David L. Buckeridge, Nigel Collier
Abstract
Infectious disease outbreaks continue to pose a significant threat to human health and well-being. To improve disease surveillance and understanding of disease spread, several surveillance systems have been developed to monitor daily news alerts and social media. However, existing systems lack thorough epidemiological analysis in relation to corresponding alerts or news, largely due to the scarcity of well-annotated reports data. To address this gap, we introduce the Biomedical Alert News Dataset (BAND), which includes 1,508 samples from existing reported news articles, open emails, and alerts, as well as 30 epidemiology-related questions. These questions necessitate the model's expert reasoning abilities, thereby offering valuable insights into the outbreak of the disease. The BAND dataset brings new challenges to the NLP world, requiring better inference capability of the content and the ability to infer important information. We provide several benchmark tasks, including Named Entity Recognition (NER), Question Answering (QA), and Event Extraction (EE), to demonstrate existing models' capabilities and limitations in handling epidemiology-specific tasks. It is worth noting that some models may lack the human-like inference capability required to fully utilize the corpus. To the best of our knowledge, the BAND corpus is the largest corpus of well-annotated biomedical outbreak alert news with elaborately designed questions, making it a valuable resource for epidemiologists and NLP researchers alike.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f9f72fc-100a-4fcb-8f64-d83a439178dcBuilds on3
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Prefix-Tuning: Optimizing Continuous Prompts for GenerationXiang Lisa Li, Percy LiangACL 2021
Related papers
- ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital EnvironmentsSourjyadip Ray, Kushal Gupta, Soumi Kundu, Payal Arvind Kasat et al.EMNLP 2024 · 1 citation
- COMETA: A Corpus for Medical Entity Linking in the Social MediaMarco Basaldella, Fangyu Liu, Ehsan Shareghi, Nigel CollierEMNLP 2020 · 4 citations
- SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLPDecheng Duan, Jitong Peng, Yingyi Zhang, Chengzhi ZhangEMNLP 2025 · 1 citation
- PHEE: A Dataset for Pharmacovigilance Event Extraction from TextZhaoyue Sun, Jiazheng Li, Gabriele Pergola, Byron C. Wallace et al.EMNLP 2022 · 17 citations
- MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive AnnotationsVishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu et al.EMNLP 2024 · 1 citation
