KinyaBERT: a Morphology-aware Kinyarwanda Language Model
Antoine Nzeyimana, Andre Niyongabo Rubungo
摘要
Pre-trained language models such as BERT have been successful at tackling many natural language processing tasks. However, the unsupervised sub-word tokenization methods commonly used in these models (e.g., byte-pair encoding -BPE) are sub-optimal at handling morphologically rich languages. Even given a morphological analyzer, naive sequencing of morphemes into a standard BERT architecture is inefficient at capturing morphological compositionality and expressing word-relative syntactic regularities. We address these challenges by proposing a simple yet effective twotier BERT architecture that leverages a morphological analyzer and explicitly represents morphological compositionality. Despite the success of BERT, most of its evaluations have been conducted on high-resource languages, obscuring its applicability on low-resource languages. We evaluate our proposed method on the low-resource morphologically rich Kinyarwanda language, naming the proposed model architecture KinyaBERT. A robust set of experimental results reveal that KinyaBERT outperforms solid baselines by 2% in F1 score on a named entity recognition task and by 4.3% in average score of a machine-translated GLUE benchmark. KinyaBERT fine-tuning has better convergence and achieves more robust results on multiple tasks even in the presence of translation noise. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio 等EMNLP 2024 · 被引用 10 次
- Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language ModelsJames A. Michaelov, Catherine Arnett, Tyler A. Chang, Ben BergenEMNLP 2023 · 被引用 6 次
- Text Rendering Strategies for Pixel Language ModelsJonas F. Lotz, Elizabeth Salesky, Phillip Rust, Desmond ElliottEMNLP 2023 · 被引用 3 次
- BERT-like Models for Slavic Morpheme SegmentationDmitry Morozov, Lizaveta Astapenka, Anna V. Glazkova, Timur Garipov 等ACL 2025 · 被引用 1 次
- Cheetah: Natural Language Generation for 517 African LanguagesIfe Adebara, AbdelRahim A. Elmadany, Muhammad Abdul-MageedACL 2024 · 被引用 1 次
它引用的顶会 Paper8
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont 等ACL 2020 · 被引用 703 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 被引用 358 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
相关 Paper
- Hints on the data for language modeling of synthetic languages with transformersRodolfo Zevallos, Núria BelACL 2023 · 被引用 2 次
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 被引用 1 次
- Effects of sub-word segmentation on performance of transformer language modelsJue Hou, Anisia Katinskaia, Anh-Duc Vu, Roman YangarberEMNLP 2023 · 被引用 2 次
- Getting The Most Out of Your Training Data: Exploring Unsupervised Tasks for Morphological InflectionAbhishek Purushothama, Adam Wiemerslage, Katharina von der WenseEMNLP 2024
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
