Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Javier Ortiz Suárez, Benoît Sagot, Abhishek Srivastava
Abstract
We introduce the first treebank for a romanized user-generated content variety of Algerian, a North-African Arabic dialect known for its frequent usage of code-switching. Made of 1500 sentences, fully annotated in morpho-syntax and Universal Dependency syntax, with full translation at both the word and the sentence levels, this treebank is made freely available. It is supplemented with 50k unlabeled sentences collected from Common Crawl and webcrawled data using intensive data-mining techniques. Preliminary experiments demonstrate its usefulness for POS tagging and dependency parsing. We believe that what we present in this paper is useful beyond the low-resource language community. This is the first time that enough unlabeled and annotated data is provided for an emerging user-generated content dialectal language with rich morphology and code switching, making it an challenging testbed for most recent NLP approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- A Golden Age: Conspiracy Theories' Relationship with Misinformation Outlets, News Media, and the Wider InternetHans W. A. Hanley, Deepak Kumar, Zakir DurumericCSCW 2023 · 22 citations
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
- TwittIrish: A Universal Dependencies Treebank of Tweets in Modern IrishLauren Cassidy, Teresa Lynn, James Barry, Jennifer FosterACL 2022
- Evaluating morphological typology in zero-shot cross-lingual transferAntonio Martínez-García, Toni Badia, Jeremy BarnesACL 2021
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road AheadJesujoba Oluwadara Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich KlakowEMNLP 2025
Builds on1
Related papers
- Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah et al.EMNLP 2024 · 5 citations
- Dialectal Coverage And Generalization in Arabic Speech RecognitionAmirbek Djanibekov, Hawau Olamide Toyin, Raghad Alshalan, Abdullah Alatir et al.ACL 2025
- ALDi: Quantifying the Arabic Level of Dialectness of TextAmr Keleg, Sharon Goldwater, Walid MagdyEMNLP 2023 · 4 citations
- Genre as Weak Supervision for Cross-lingual Dependency ParsingMax Müller-Eberstein, Rob van der Goot, Barbara PlankEMNLP 2021 · 13 citations
- The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation ProjectAngelina Aspra Aquino, Lester James Validad Miranda, Elsie Marie T. OrACL 2025
