UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Reynolds
摘要
We introduce UNIVERSALCEFR, a largescale multilingual and multidimensional dataset of texts annotated with CEFR (Common European Framework of Reference) levels in 13 languages. To enable open research in automated readability and language proficiency assessment, UNIVERSALCEFR comprises 505,807 CEFR-labeled texts curated from educational and learner-oriented resources, standardized into a unified data format to support consistent processing, analysis, and modelling across tasks and languages. To demonstrate its utility, we conduct benchmarking experiments using three modelling paradigms: a) linguistic feature-based classification, b) fine-tuning pre-trained LLMs, and c) descriptor-based prompting of instruction-tuned LLMs. Our results support using linguistic features and fine-tuning pretrained models in multilingual CEFR level assessment. Overall, UNIVER-SALCEFR aims to establish best practices in data distribution for language proficiency research by standardising dataset formats, and promoting their accessibility to the global research community. universalcefr.github.io huggingface.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- CEFR-Based Sentence Difficulty Annotation and AssessmentYuki Arase, Satoru Uchida, Tomoyuki KajiwaraEMNLP 2022 · 被引用 17 次
- DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document SimplificationRegina Stodden, Omar Momen, Laura KallmeyerACL 2023 · 被引用 3 次
相关 Paper
- ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability AssessmentTarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra 等EMNLP 2024 · 被引用 7 次
- OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ LanguagesChester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 等EMNLP 2025 · 被引用 2 次
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza 等ICLR 2026 · 被引用 22 次
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera 等ACL 2026 · 被引用 5 次
- Targeted Syntactic Evaluation for Grammatical Error CorrectionAomi Koyama, Masato Mita, Su-Youn Yoon, Yasufumi Takama 等ACL 2025
