IDRISI-RA: The First Arabic Location Mention Recognition Dataset of Disaster Tweets
Reem Suwaileh, Muhammad Imran, Tamer Elsayed
Abstract
Extracting geolocation information from social media data enables effective disaster management, as it helps response authorities; for example, in locating incidents for planning rescue activities, and affected people for evacuation. Nevertheless, geolocation extraction is greatly understudied for the low resource languages such as Arabic. To fill this gap, we introduce IDRISI-RA, the first publicly-available Arabic Location Mention Recognition (LMR) dataset that provides human- and automatically-labeled versions in order of thousands and millions of tweets, respectively. It contains both location mentions and their types (e.g., district, city). Our extensive analysis shows the decent geographical, domain, location granularity, temporal, and dialectical coverage of IDRISI-RA. Furthermore, we establish baselines using the standard Arabic NER models and build two simple, yet effective, LMR models. Our rigorous experiments confirm the need for developing specific models for Arabic LMR in the disaster domain. Moreover, experiments show the promising domain and geographical generalizability of IDRISI-RA under zero-shot learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- #Outage: Detecting Power and Communication Outages from Social NetworksUdit Paul, Alexander Ermakov, Michael Nekrasov, Vivek Adarsh et al.WWW 2020 · 24 citations
- AraVQA: Building a New Arabic Factoid Visual Question Answering Dataset from WikipediaSultan Alrowili, Younes Samih, Abed Alhakim Freihat, Mathan Kumar EswaranACL 2026
- Detecting Perceived Emotions in Hurricane DisastersShrey Desai, Cornelia Caragea, Junyi Jessy LiACL 2020 · 2 citations
- SPEED++: A Multilingual Event Extraction Framework for Epidemic Prediction and PreparednessTanmay Parekh, Jeffrey Kwan, Jiarui Yu, Sparsh Johri et al.EMNLP 2024 · 4 citations
- Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsAbdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi et al.ACL 2026
