Unsupervised Readability Assessment via Learning from Weak Readability Signals
Yuliang Liu, Zhiwei Jiang, Yafeng Yin, Cong Wang, Sheng Chen, Zhaoling Chen, Qing Gu
摘要
Unsupervised readability assessment aims to evaluate the reading difficulty of text without any manually-labeled data for model training. This is a challenging task because the absence of labeled data makes it difficult for the model to understand what readability is. In this paper, we propose a novel framework to Learn a neural model from Weak Readability Signals (LWRS). Instead of relying on labeled data, LWRS utilizes a set of heuristic signals that specialize in describing text readability from different aspects to guide the model in outputting readability scores for ranking. Specifically, to effectively use multiple heuristic weak signals for model training, we build a multi-signal learning model that ranks the unlabeled texts from multiple readability-related aspects based on intra- and inter-signal learning. We also adopt the pairwise ranking paradigm to reduce the cascade coupling among partial-order pairs. Furthermore, we propose identifying the most representative signal based on the batch-level consensus distribution of all signals. This strategy helps identify the predicted signal that is most correlated with readability in the absence of ground-truth labels. We conduct experiments on three public readability assessment datasets. The experimental results demonstrate that our LWRS outperforms each heuristic signal and their combinations significantly, and can even perform comparably with some supervised methods. Additionally, our LWRS trained on one dataset can be effectively transferred to other datasets, including those in other languages, which indicates its good generalization and potential for wide application.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic FeaturesBruce W. Lee, Yoo Sung Jang, Jason Hyung-Jong LeeEMNLP 2021 · 被引用 46 次
- Learning Syntactic Dense Embedding with Correlation Graph for Automatic Readability AssessmentXinying Qiu, Yuan Chen, Hanwu Chen, Jian-Yun Nie 等ACL 2021
相关 Paper
- Aggregating Multiple Heuristic Signals as Supervision for Unsupervised Automated Essay ScoringCong Wang, Zhiwei Jiang, Yafeng Yin, Zifeng Cheng 等ACL 2023 · 被引用 4 次
- Zero-shot Large Language Models for Automatic Readability AssessmentRiley Grossman, Yi ChenACL 2026 · 被引用 1 次
- A Unified Neural Network Model for Readability Assessment with Feature Projection and Length-Balanced LossWenbiao Li, Ziyang Wang, Yunfang WuEMNLP 2022 · 被引用 7 次
- Generalizable Video Quality Assessment via Weak-to-Strong LearningLinhan Cao, Wei Sun, Xiangyang Zhu, Kaiwei Zhang 等CVPR 2026 · 被引用 9 次
- Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMsYu-Wen Chen, Melody Ma, Julia HirschbergEMNLP 2025
