Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features
Bruce W. Lee, Yoo Sung Jang, Jason Hyung-Jong Lee
摘要
We report two essential improvements in readability assessment: 1. three novel features in advanced semantics and 2. the timely evidence that traditional ML models (e.g. Random Forest, using handcrafted features) can combine with transformers (e.g. RoBERTa) to augment model performance. First, we explore suitable transformers and traditional ML models. Then, we extract 255 handcrafted linguistic features using self-developed extraction software. Finally, we assemble those to create several hybrid models, achieving state-of-the-art (SOTA) accuracy on popular datasets in readability assessment. The use of handcrafted features help model performance on smaller datasets. Notably, our RoBERTA-RF-T1 hybrid achieves the near-perfect classification accuracy of 99%, a 20.3% increase from the previous SOTA. Table 11: Entity Density Features (EnDF).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- A Unified Neural Network Model for Readability Assessment with Feature Projection and Length-Balanced LossWenbiao Li, Ziyang Wang, Yunfang WuEMNLP 2022 · 被引用 7 次
- Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRAMaharshi Gor, Hal Daumé III, Tianyi Zhou, Jordan L. Boyd-GraberEMNLP 2024 · 被引用 7 次
- Unsupervised Readability Assessment via Learning from Weak Readability SignalsYuliang Liu, Zhiwei Jiang, Yafeng Yin, Cong Wang 等SIGIR 2023 · 被引用 3 次
- MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical DomainChao Jiang, Wei XuEMNLP 2024 · 被引用 3 次
- InterpretARA: Enhancing Hybrid Automatic Readability Assessment with Linguistic Feature Interpreter and Contrastive LearningJinshan Zeng, Xianchao Tong, Xianglong Yu, Wenyan Xiao 等AAAI 2024 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Learning Syntactic Dense Embedding with Correlation Graph for Automatic Readability AssessmentXinying Qiu, Yuan Chen, Hanwu Chen, Jian-Yun Nie 等ACL 2021
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Self-Supervised Collaborative Information Bottleneck for Text Readability AssessmentJinshan Zeng, Xianglong Yu, Xianchao Tong, Wenyan XiaoAAAI 2025
- Syntax-Enhanced Pre-trained ModelZenan Xu, Daya Guo, Duyu Tang, Qinliang Su 等ACL 2021
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
