Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks
Felermino Dário Mário António Ali, Henrique Lopes Cardoso, Rui Sousa-Silva
Abstract
This paper introduces a comprehensive collection of NLP resources for Emakhuwa, Mozambique's most widely spoken language. The resources include the first manually translated news bitext corpus between Portuguese and Emakhuwa, news topic classification datasets, and monolingual data. We detail the process and challenges of acquiring this data and present benchmark results for machine translation and news topic classification tasks. Our evaluation examines the impact of different data types-originally clean text, postcorrected OCR, and back-translated data-and the effects of fine-tuning from pre-trained models, including those focused on African languages. Our benchmarks demonstrate good performance in news topic classification and promising results in machine translation. We fine-tuned multilingual encoder-decoder models using real and synthetic data and evaluated them on our test set and the FLORES evaluation sets. The results highlight the importance of incorporating more data and potential for future improvements. All models, code, and datasets are available in the https://huggingface.co/LIACC repository under the CC BY 4.0 license.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55e10a71-4839-45a2-bb2c-73617cf2f312Cited by top-tier papers2
- BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle et al.ACL 2025 · 81 citations
- Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual ContextFelermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-SilvaEMNLP 2025
Builds on5
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Rethinking Embedding Coupling in Pre-trained Language ModelsHyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson et al.ICLR 2021 · 11 citations
- OCR Post Correction for Endangered Language TextsShruti Rijhwani, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 1 citation
- Cheetah: Natural Language Generation for 517 African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-MageedACL 2024 · 1 citation
Related papers
- AFRIDOC-MT: Document-level MT Corpus for African LanguagesJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet et al.EMNLP 2025
- AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African LanguagesMachel Reid, Junjie Hu, Graham Neubig, Yutaka MatsuoEMNLP 2021 · 12 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLPSheriff Issaka, Keyi Wang, Yinka Ajibola, Oluwatumininu Samuel-Ipaye et al.ACL 2026
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesPedro Javier Ortiz Suárez, Laurent Romary, Benoît SagotACL 2020 · 72 citations
