Unsupervised Readability Assessment via Learning from Weak Readability Signals
Yuliang Liu, Zhiwei Jiang, Yafeng Yin, Cong Wang, Sheng Chen, Zhaoling Chen, Qing Gu
Abstract
Unsupervised readability assessment aims to evaluate the reading difficulty of text without any manually-labeled data for model training. This is a challenging task because the absence of labeled data makes it difficult for the model to understand what readability is. In this paper, we propose a novel framework to Learn a neural model from Weak Readability Signals (LWRS). Instead of relying on labeled data, LWRS utilizes a set of heuristic signals that specialize in describing text readability from different aspects to guide the model in outputting readability scores for ranking. Specifically, to effectively use multiple heuristic weak signals for model training, we build a multi-signal learning model that ranks the unlabeled texts from multiple readability-related aspects based on intra- and inter-signal learning. We also adopt the pairwise ranking paradigm to reduce the cascade coupling among partial-order pairs. Furthermore, we propose identifying the most representative signal based on the batch-level consensus distribution of all signals. This strategy helps identify the predicted signal that is most correlated with readability in the absence of ground-truth labels. We conduct experiments on three public readability assessment datasets. The experimental results demonstrate that our LWRS outperforms each heuristic signal and their combinations significantly, and can even perform comparably with some supervised methods. Additionally, our LWRS trained on one dataset can be effectively transferred to other datasets, including those in other languages, which indicates its good generalization and potential for wide application.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c05af9cb-1a4a-408d-920a-a9242a260c5dBuilds on2
- Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic FeaturesBruce W. Lee, Yoo Sung Jang, Jason Hyung-Jong LeeEMNLP 2021 · 46 citations
- Learning Syntactic Dense Embedding with Correlation Graph for Automatic Readability AssessmentXinying Qiu, Yuan Chen, Hanwu Chen, Jian-Yun Nie et al.ACL 2021
Related papers
- Aggregating Multiple Heuristic Signals as Supervision for Unsupervised Automated Essay ScoringCong Wang, Zhiwei Jiang, Yafeng Yin, Zifeng Cheng et al.ACL 2023 · 4 citations
- Zero-shot Large Language Models for Automatic Readability AssessmentRiley Grossman, Yi ChenACL 2026 · 1 citation
- A Unified Neural Network Model for Readability Assessment with Feature Projection and Length-Balanced LossWenbiao Li, Ziyang Wang, Yunfang WuEMNLP 2022 · 7 citations
- Generalizable Video Quality Assessment via Weak-to-Strong LearningLinhan Cao, Wei Sun, Xiangyang Zhu, Kaiwei Zhang et al.CVPR 2026 · 9 citations
- Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMsYu-Wen Chen, Melody Ma, Julia HirschbergEMNLP 2025
