TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
Ezgi Basar, Francesca Padovani, Jaap Jumelet, Arianna Bisazza
摘要
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 被引用 5 次
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 被引用 3 次
它引用的顶会 Paper3
- SLING: Sino Linguistic Evaluation of Large Language ModelsYixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit IyyerEMNLP 2022 · 被引用 7 次
- RuBLiMP: Russian Benchmark of Linguistic Minimal PairsEkaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova 等EMNLP 2024 · 被引用 2 次
- Cross-Linguistic Syntactic Evaluation of Word Prediction ModelsAaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou, Natalia Talmina 等ACL 2020 · 被引用 2 次
相关 Paper
- CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language ModelsMiyu Oba, Saku SugawaraACL 2026
- TUMLU: A Unified and Native Language Understanding Benchmark for Turkic LanguagesJafar Isbarov, Arofat Akhundjanova, Mammad Hajili, Kavsar Huseynova 等ACL 2025
- Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelLeonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai 等EMNLP 2023 · 被引用 10 次
- Data and Representation for Turkish Natural Language InferenceEmrah Budur, Riza Özçelik, Tunga Gungor, Christopher PottsEMNLP 2020
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
