UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds
Abstract
We introduce UNIVERSALCEFR, a largescale multilingual and multidimensional dataset of texts annotated with CEFR (Common European Framework of Reference) levels in 13 languages. To enable open research in automated readability and language proficiency assessment, UNIVERSALCEFR comprises 505,807 CEFR-labeled texts curated from educational and learner-oriented resources, standardized into a unified data format to support consistent processing, analysis, and modelling across tasks and languages. To demonstrate its utility, we conduct benchmarking experiments using three modelling paradigms: a) linguistic feature-based classification, b) fine-tuning pre-trained LLMs, and c) descriptor-based prompting of instruction-tuned LLMs. Our results support using linguistic features and fine-tuning pretrained models in multilingual CEFR level assessment. Overall, UNIVER-SALCEFR aims to establish best practices in data distribution for language proficiency research by standardising dataset formats, and promoting their accessibility to the global research community. universalcefr.github.io huggingface.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- CEFR-Based Sentence Difficulty Annotation and AssessmentYuki Arase, Satoru Uchida, Tomoyuki KajiwaraEMNLP 2022 · 17 citations
- DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document SimplificationRegina Stodden, Omar Momen, Laura KallmeyerACL 2023 · 3 citations
Related papers
- ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability AssessmentTarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra et al.EMNLP 2024 · 7 citations
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä et al.EMNLP 2025 · 2 citations
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza et al.ICLR 2026 · 22 citations
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera et al.ACL 2026 · 5 citations
- Targeted Syntactic Evaluation for Grammatical Error CorrectionAomi Koyama, Masato Mita, Su-Youn Yoon, Yasufumi Takama et al.ACL 2025
