SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and Preparedness
Tanmay Parekh, Jeffrey Kwan, Jiarui Yu, Sparsh Johri, Hyosang Ahn, Sreya Muppalla, Kai-Wei Chang, Wei Wang, Nanyun Peng
Abstract
Social media is often the first place where communities discuss the latest societal trends. Prior works have utilized this platform to extract epidemic-related information (e.g. infections, preventive measures) to provide early warnings for epidemic prediction. However, these works only focused on English posts, while epidemics can occur anywhere in the world, and early discussions are often in the local, non-English languages. In this work, we introduce the first multilingual Event Extraction (EE) framework SPEED++ for extracting epidemic event information for a wide range of diseases and languages. To this end, we extend a previous epidemic ontology with 20 argument roles; and curate our multilingual EE dataset SPEED++ comprising 5.1K tweets in four languages for four diseases. Annotating data in every language is infeasible; thus we develop zero-shot cross-lingual cross-disease models (i.e., training only on English COVID data) utilizing multilingual pre-training and show their efficacy in extracting epidemic-related events for 65 diverse languages across different diseases. Experiments demonstrate that our framework can provide epidemic warnings for COVID-19 in its earliest stages in Dec 2019 (3 weeks before global discussions) from Chinese Weibo posts without any training in Chinese. Furthermore, we exploit our framework's argument extraction capabilities to aggregate community epidemic discussions like symptoms and cure measures, aiding misinformation detection and public attention monitoring. Overall, we lay a strong foundation for multilingual epidemic preparedness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c3da6af-17c4-4bd1-bd22-1576b7ecff9aCited by top-tier papers2
- SNaRe: Domain-aware Data Generation for Low-Resource Event DetectionTanmay Parekh, Yuxuan Dong, Lucas Bandarkar, Artin Kim et al.EMNLP 2025 · 1 citation
- DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM ReasoningTanmay Parekh, Kartik Mehta, Ninareh Mehrabi, Kai-Wei Chang et al.EMNLP 2025 · 1 citation
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Event Extraction by Answering (Almost) Natural QuestionsXinya Du, Claire CardieEMNLP 2020 · 391 citations
- A Joint Neural Model for Information Extraction with Global FeaturesYing Lin, Heng Ji, Fei Huang, Lingfei WuACL 2020 · 376 citations
- MAVEN: A Massive General Domain Event Detection DatasetXiaozhi Wang, Ziqi Wang, Xu Han, Wangyi Jiang et al.EMNLP 2020 · 143 citations
Related papers
- MEE: A Novel Multilingual Event Extraction DatasetAmir Pouran Ben Veyseh, Javid Ebrahimi, Franck Dernoncourt, Thien Huu NguyenEMNLP 2022 · 3 citations
- Zero-Shot Rumor Detection with Propagation Structure via Prompt LearningHongzhan Lin, Pengyao Yi, Jing Ma, Haiyun Jiang et al.AAAI 2023 · 84 citations
- Learning from Tweets: Opportunities and Challenges to Inform Policy Making During Dengue EpidemicFarhana Shahid, Shahinul Hoque Ony, Takrim Rahman Albi, Sriram Chellappan et al.CSCW 2020 · 28 citations
- COVID-19 Vaccine Misinformation in Middle Income CountriesJongin Kim, Byeo Bak, Aditya Agrawal, Jiaxi Wu et al.EMNLP 2023 · 3 citations
- Multilingual Generative Language Models for Zero-Shot Cross-Lingual Event Argument ExtractionKuan-Hao Huang, I-Hung Hsu, Prem Natarajan, Kai-Wei Chang et al.ACL 2022
