RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
Ekaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova, Ekaterina Artemova, Vladislav Mikhailov
摘要
Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. This paper introduces the Russian Benchmark of Linguistic Minimal Pairs (RuBLiMP), which includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text corpora and decontaminating test data. We describe the data collection protocol and present the results of evaluating 25 language models in various scenarios. We find that the widely used LMs for Russian are sensitive to morphological and agreement-oriented contrasts, but fall behind humans on phenomena requiring the understanding of structural relations, negation, transitivity, and tense. RuBLiMP, the codebase, and other materials are publicly available. * Equal contribution. † Work is partially done while at HSE University. Parse Sentence Extraction (a) root spal 'sleep' Vpervye 'For the first time' kosmonavt 'an astronaut' nevesomosti 'zero gravity' v 'in' obl case nsubj advmod Sentence Annotation Minimal Pair Generation adposition government Vpervye kosmonavt spal v nevesomosti. Vpervye kosmonavt spal v nevesomost'. subject-predicate agreement (number) Vpervye kosmonavt spal v nevesomosti. Vpervye kosmonavt spali v nevesomosti. subject-predicate agreement (gender) Vpervye kosmonavt spal v nevesomosti. Vpervye kosmonavt spalo v nevesomosti.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Minimal Pair-Based Evaluation of Code-SwitchingIgor Sterner, Simone TeufelACL 2025 · 被引用 8 次
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 被引用 5 次
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 被引用 3 次
- TurBLiMP: A Turkish Benchmark of Linguistic Minimal PairsEzgi Basar, Francesca Padovani, Jaap Jumelet, Arianna BisazzaEMNLP 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
相关 Paper
- RuCoLA: Russian Corpus of Linguistic AcceptabilityVladislav Mikhailov, Tatiana Shamardina, Max Ryabinin, Alena Pestova 等EMNLP 2022 · 被引用 19 次
- SLING: Sino Linguistic Evaluation of Large Language ModelsYixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit IyyerEMNLP 2022 · 被引用 7 次
- RussianSuperGLUE: A Russian Language Understanding Evaluation BenchmarkTatiana Shavrina, Alena Fenogenova, Anton A. Emelyanov, Denis Shevelev 等EMNLP 2020 · 被引用 11 次
- CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language ModelsMiyu Oba, Saku SugawaraACL 2026
- Cross-Linguistic Syntactic Evaluation of Word Prediction ModelsAaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou, Natalia Talmina 等ACL 2020 · 被引用 2 次
