TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
Ezgi Basar, Francesca Padovani, Jaap Jumelet, Arianna Bisazza
Abstract
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de9164c2-d0f1-4213-ba91-0494ee21263eCited by top-tier papers2
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 5 citations
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 3 citations
Builds on3
- SLING: Sino Linguistic Evaluation of Large Language ModelsYixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit IyyerEMNLP 2022 · 7 citations
- RuBLiMP: Russian Benchmark of Linguistic Minimal PairsEkaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova et al.EMNLP 2024 · 2 citations
- Cross-Linguistic Syntactic Evaluation of Word Prediction ModelsAaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou, Natalia Talmina et al.ACL 2020 · 2 citations
Related papers
- CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language ModelsMiyu Oba, Saku SugawaraACL 2026
- TUMLU: A Unified and Native Language Understanding Benchmark for Turkic LanguagesJafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova et al.ACL 2025
- Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelLeonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai et al.EMNLP 2023 · 10 citations
- Data and Representation for Turkish Natural Language InferenceEmrah Budur, Riza Özçelik, Tunga Gungor, Christopher PottsEMNLP 2020
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
