Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks
Felermino Dário Mário António Ali, Henrique Lopes Cardoso, Rui Sousa-Silva
摘要
This paper introduces a comprehensive collection of NLP resources for Emakhuwa, Mozambique's most widely spoken language. The resources include the first manually translated news bitext corpus between Portuguese and Emakhuwa, news topic classification datasets, and monolingual data. We detail the process and challenges of acquiring this data and present benchmark results for machine translation and news topic classification tasks. Our evaluation examines the impact of different data types-originally clean text, postcorrected OCR, and back-translated data-and the effects of fine-tuning from pre-trained models, including those focused on African languages. Our benchmarks demonstrate good performance in news topic classification and promising results in machine translation. We fine-tuned multilingual encoder-decoder models using real and synthetic data and evaluated them on our test set and the FLORES evaluation sets. The results highlight the importance of incorporating more data and potential for future improvements. All models, code, and datasets are available in the https://huggingface.co/LIACC repository under the CC BY 4.0 license.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle 等ACL 2025 · 被引用 81 次
- Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual ContextFelermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-SilvaEMNLP 2025
它引用的顶会 Paper5
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Rethinking Embedding Coupling in Pre-trained Language ModelsHyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson 等ICLR 2021 · 被引用 11 次
- OCR Post Correction for Endangered Language TextsShruti Rijhwani, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 被引用 1 次
- Cheetah: Natural Language Generation for 517 African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-MageedACL 2024 · 被引用 1 次
相关 Paper
- AFRIDOC-MT: Document-level MT Corpus for African LanguagesJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet 等EMNLP 2025
- AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African LanguagesMachel Reid, Junjie Hu, Graham Neubig, Yutaka MatsuoEMNLP 2021 · 被引用 12 次
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe 等ACL 2024 · 被引用 30 次
- The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLPSheriff Issaka, Keyi Wang, Yinka Ajibola, Oluwatumininu Samuel-Ipaye 等ACL 2026
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesPedro Javier Ortiz Suárez, Laurent Romary, Benoît SagotACL 2020 · 被引用 72 次
