ACL2026

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

Farah Adeeba, Brian Dillon, Hassan Sajjad, Rajesh Bhatt

2 citations

Abstract

Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages; however, they often include significantly less data for low-resource languages such as Urdu compared to highresource languages like English. To assess the linguistic knowledge of LLMs in Urdu, we present the Urdu Benchmark of Linguistic Minimal Pairs (UrBLiMP) i.e. pairs of minimally different sentences that contrast in grammatical acceptability. UrBLiMP comprises 5,696 minimal pairs targeting ten core syntactic phenomena, carefully curated using the Urdu Treebank and diverse Urdu text corpora. A human evaluation of UrBLiMP annotations yielded a 96.10% inter-annotator agreement, confirming the reliability of the dataset. We evaluate twenty multilingual LLMs on UrBLiMP, revealing significant variation in performance across linguistic phenomena. While LLaMA-3-70B achieves the highest average accuracy (94.73%), its performance is statistically comparable to other top models such as Gemma-3-27B-PT. These findings highlight both the potential and the limitations of current multilingual LLMs in capturing fine-grained syntactic knowledge in low-resource languages. Language Size P Method BLiMP English 67K 67 Dict & Templates CLiMP Chinese 16K 16 Translation & Templates SLING Chinese 38K 38 UD & Templates TurBLiMP Turkish 16K 16 Semi-automatically RuBLiMP Russian 45K 45 Semi-automatically JBLIMP Japanese 331 39 Extracted from articles BLiMP-NL Dutch 8.4K 84 Semi-automatically MultiBLiMP 101 Languages 128K 2 UD BHS Basque, Hindi, Swahili 300 3 Dictionary & Templates UrBLiMP Urdu 5696 19 Text corpus, rules C Use of AI-Assistant ChatGPT was used to proofread and improve the text of this paper by correcting grammatical, spelling, and stylistic errors.