ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
Tarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra, Wei Xu
Abstract
We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This paper introduces README++, a multilingual multi-domain dataset with human annotations of 9757 sentences in Arabic, English, French, Hindi, and Russian, collected from 112 different data sources. This benchmark will encourage research on developing robust multilingual readability assessment methods. Using README++, we benchmark multilingual and monolingual language models in the supervised, unsupervised, and few-shot prompting settings. The domain and language diversity in README++ enable us to test more effective few-shot prompting, and identify shortcomings in state-of-the-art unsupervised methods. Our experiments also reveal exciting results of superior domain generalization and enhanced cross-lingual transfer capabilities by models trained on README++. We will make our data publicly available and release a python package tool for multilingual sentence readability prediction using our trained models at: https://github.com/ tareknaous/readme Dataset Languages Scripts #Data Sources
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Zero-shot Large Language Models for Automatic Readability AssessmentRiley Grossman, Yi ChenACL 2026 · 1 citation
- Assessing French Readability for Adults with Low Literacy: A Global and Local PerspectiveWafa Aissa, Thibault Bañeras-Roux, Elodie Vanzeveren, Lingyun Gao et al.EMNLP 2025
- Question Difficulty Estimation for Large Language Models via Answer Plausibility ScoringJamshid Mozafari, Bhawna Piryani, Adam JatowtACL 2026
- Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text SimplificationJinhong Jeong, Junghun Park, Youngjae YuACL 2026
Builds on10
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Simple or Complex? Learning to Predict Readability of Bengali TextsSusmoy Chakraborty, Mir Tafseer Nayeem, Wasi Uddin AhmadAAAI 2021 · 21 citations
- CEFR-Based Sentence Difficulty Annotation and AssessmentYuki Arase, Satoru Uchida, Tomoyuki KajiwaraEMNLP 2022 · 17 citations
- MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual CaptioningBang Yang, Fenglin Liu, Xian Wu, Yaowei Wang et al.ACL 2023 · 10 citations
Related papers
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical DomainChao Jiang, Wei XuEMNLP 2024 · 3 citations
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkDominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas et al.EMNLP 2023 · 25 citations
- Revisiting non-English Text Simplification: A Unified Multilingual BenchmarkMichael J. Ryan, Tarek Naous, Wei XuACL 2023 · 14 citations
- P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMsYidan Zhang, Yu Wan, Boyi Deng, Baosong Yang et al.EMNLP 2025
