RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark
Tatiana Shavrina, Alena Fenogenova, Anton A. Emelyanov, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, Andrey Evlampiev
Abstract
In this paper, we introduce an advanced Russian general language understanding evaluation benchmark -RussianGLUE. Recent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills -detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology (Wang et al., 2019), was developed from scratch for the Russian language. We provide baselines, human level evaluation, an opensource framework for evaluating models and an overall leaderboard of transformer models for the Russian language. Besides, we present the first results of comparing multilingual models in the adapted diagnostic test set and offer the first steps to further expanding or assessing state-of-the-art models independently of language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f987df61-fa0f-486f-b570-ba0116df495cCited by top-tier papers12
- CBLUE: A Chinese Biomedical Language Understanding Evaluation BenchmarkNingyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang et al.ACL 2022 · 242 citations
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesEmanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy et al.ICML 2022 · 71 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid et al.EMNLP 2022 · 17 citations
- MERA: A Comprehensive LLM Evaluation in RussianAlena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova et al.ACL 2024 · 10 citations
Builds on1
Related papers
- bgGLUE: A Bulgarian General Language Understanding Evaluation BenchmarkMomchil Hardalov, Pepa Atanasova, Todor Mihaylov, Galia Angelova et al.ACL 2023 · 4 citations
- BelarusianGLUE: Towards a Natural Language Understanding Benchmark for BelarusianMaksim Aparovich, Volha Harytskaya, Vladislav Poritski, Oksana Volchek et al.ACL 2025
- Superlim: A Swedish Language Understanding Evaluation BenchmarkAleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger et al.EMNLP 2023 · 2 citations
- AwarenessBench: Assessing Cognitive Capabilities of Language ModelsXiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang et al.ACL 2026
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II et al.ACL 2022
