Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features
Bruce W. Lee, Yoo Sung Jang, Jason Hyung-Jong Lee
Abstract
We report two essential improvements in readability assessment: 1. three novel features in advanced semantics and 2. the timely evidence that traditional ML models (e.g. Random Forest, using handcrafted features) can combine with transformers (e.g. RoBERTa) to augment model performance. First, we explore suitable transformers and traditional ML models. Then, we extract 255 handcrafted linguistic features using self-developed extraction software. Finally, we assemble those to create several hybrid models, achieving state-of-the-art (SOTA) accuracy on popular datasets in readability assessment. The use of handcrafted features help model performance on smaller datasets. Notably, our RoBERTA-RF-T1 hybrid achieves the near-perfect classification accuracy of 99%, a 20.3% increase from the previous SOTA. Table 11: Entity Density Features (EnDF).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50376159-a9a0-4072-a672-b6a1f76e4e24Cited by top-tier papers10
- A Unified Neural Network Model for Readability Assessment with Feature Projection and Length-Balanced LossWenbiao Li, Ziyang Wang, Yunfang WuEMNLP 2022 · 7 citations
- Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRAMaharshi Gor, Hal Daumé III, Tianyi Zhou, Jordan L. Boyd-GraberEMNLP 2024 · 7 citations
- Unsupervised Readability Assessment via Learning from Weak Readability SignalsYuliang Liu, Zhiwei Jiang, Yafeng Yin, Cong Wang et al.SIGIR 2023 · 3 citations
- MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical DomainChao Jiang, Wei XuEMNLP 2024 · 3 citations
- InterpretARA: Enhancing Hybrid Automatic Readability Assessment with Linguistic Feature Interpreter and Contrastive LearningJinshan Zeng, Xianchao Tong, Xianglong Yu, Wenyan Xiao et al.AAAI 2024 · 3 citations
Builds on1
Related papers
- Learning Syntactic Dense Embedding with Correlation Graph for Automatic Readability AssessmentXinying Qiu, Yuan Chen, Hanwu Chen, Jian-Yun Nie et al.ACL 2021
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Self-Supervised Collaborative Information Bottleneck for Text Readability AssessmentJinshan Zeng, Xianglong Yu, Xianchao Tong, Wenyan XiaoAAAI 2025
- Syntax-Enhanced Pre-trained ModelZenan Xu, Daya Guo, Duyu Tang, Qinliang Su et al.ACL 2021
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
