Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Javier Ortiz Suárez, Benoît Sagot, Abhishek Srivastava
摘要
We introduce the first treebank for a romanized user-generated content variety of Algerian, a North-African Arabic dialect known for its frequent usage of code-switching. Made of 1500 sentences, fully annotated in morpho-syntax and Universal Dependency syntax, with full translation at both the word and the sentence levels, this treebank is made freely available. It is supplemented with 50k unlabeled sentences collected from Common Crawl and webcrawled data using intensive data-mining techniques. Preliminary experiments demonstrate its usefulness for POS tagging and dependency parsing. We believe that what we present in this paper is useful beyond the low-resource language community. This is the first time that enough unlabeled and annotated data is provided for an emerging user-generated content dialectal language with rich morphology and code switching, making it an challenging testbed for most recent NLP approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Golden Age: Conspiracy Theories' Relationship with Misinformation Outlets, News Media, and the Wider InternetHans W. A. Hanley, Deepak Kumar, Zakir DurumericCSCW 2023 · 被引用 22 次
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 等EMNLP 2025 · 被引用 2 次
- TwittIrish: A Universal Dependencies Treebank of Tweets in Modern IrishLauren Cassidy, Teresa Lynn, James Barry, Jennifer FosterACL 2022
- Evaluating morphological typology in zero-shot cross-lingual transferAntonio Martínez-García, Toni Badia, Jeremy BarnesACL 2021
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road AheadJesujoba Oluwadara Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich KlakowEMNLP 2025
它引用的顶会 Paper1
相关 Paper
- Casablanca: Data and Models for Multidialectal Arabic Speech RecognitionBashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah 等EMNLP 2024 · 被引用 5 次
- Dialectal Coverage And Generalization in Arabic Speech RecognitionAmirbek Djanibekov, Hawau Olamide Toyin, Raghad Alshalan, Abdullah Alatir 等ACL 2025
- ALDi: Quantifying the Arabic Level of Dialectness of TextAmr Keleg, Sharon Goldwater, Walid MagdyEMNLP 2023 · 被引用 4 次
- Genre as Weak Supervision for Cross-lingual Dependency ParsingMax Müller-Eberstein, Rob van der Goot, Barbara PlankEMNLP 2021 · 被引用 13 次
- The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation ProjectAngelina Aspra Aquino, Lester James Validad Miranda, Elsie Marie T. OrACL 2025
