RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark
Tatiana Shavrina, Alena Fenogenova, Anton A. Emelyanov, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, Andrey Evlampiev
摘要
In this paper, we introduce an advanced Russian general language understanding evaluation benchmark -RussianGLUE. Recent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills -detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology (Wang et al., 2019), was developed from scratch for the Russian language. We provide baselines, human level evaluation, an opensource framework for evaluating models and an overall leaderboard of transformer models for the Russian language. Besides, we present the first results of comparing multilingual models in the adapted diagnostic test set and offer the first steps to further expanding or assessing state-of-the-art models independently of language.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- CBLUE: A Chinese Biomedical Language Understanding Evaluation BenchmarkNingyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang 等ACL 2022 · 被引用 242 次
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesEmanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy 等ICML 2022 · 被引用 71 次
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe 等ACL 2024 · 被引用 30 次
- Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid 等EMNLP 2022 · 被引用 17 次
- MERA: A Comprehensive LLM Evaluation in RussianAlena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova 等ACL 2024 · 被引用 10 次
它引用的顶会 Paper1
相关 Paper
- bgGLUE: A Bulgarian General Language Understanding Evaluation BenchmarkMomchil Hardalov, Pepa Atanasova, Todor Mihaylov, Galia Angelova 等ACL 2023 · 被引用 4 次
- BelarusianGLUE: Towards a Natural Language Understanding Benchmark for BelarusianMaksim Aparovich, Volha Harytskaya, Vladislav Poritski, Oksana Volchek 等ACL 2025
- Superlim: A Swedish Language Understanding Evaluation BenchmarkAleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger 等EMNLP 2023 · 被引用 2 次
- AwarenessBench: Assessing Cognitive Capabilities of Language ModelsXiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang 等ACL 2026
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II 等ACL 2022
