MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical Domain
Chao Jiang, Wei Xu
摘要
Medical texts are notoriously challenging to read. Properly measuring their readability is the first step towards making them more accessible. In this paper, we present a systematic study on fine-grained readability measurements in the medical domain at both sentence-level and span-level. We introduce a new dataset MedReadMe, which consists of manually annotated readability ratings and fine-grained complex span annotation for 4,520 sentences, featuring two novel "Google-Easy" and "Google-Hard" categories. It supports our quantitative analysis, which covers 650 linguistic features and automatic complex word and jargon identification. Enabled by our high-quality annotation, we benchmark and improve several state-of-the-art sentence-level readability metrics for the medical domain specifically, which include unsupervised, supervised, and prompting-based methods using recently developed large language models (LLMs). Informed by our fine-grained complex span annotation, we find that adding a single feature, capturing the number of jargon spans, into existing readability formulas can significantly improve their correlation with human judgments. We will publicly release the dataset and code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
- Expertise Style Transfer: A New Task Towards Better Communication between Experts and LaymenYixin Cao, Ruihao Shui, Liangming Pan, Min-Yen Kan 等ACL 2020 · 被引用 50 次
- Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic FeaturesBruce W. Lee, Yoo Sung Jang, Jason Hyung-Jong LeeEMNLP 2021 · 被引用 46 次
- Making Science Simple: Corpora for the Lay Summarisation of Scientific LiteratureTomas Goldsack, Zhihao Zhang, Chenghua Lin, Carolina ScartonEMNLP 2022 · 被引用 38 次
- LENS: A Learnable Evaluation Metric for Text SimplificationMounica Maddela, Yao Dou, David Heineman, Wei XuACL 2023 · 被引用 21 次
相关 Paper
- MedJEx: A Medical Jargon Extraction Model with Wiki's Hyperlink Span and Contextualized Masked Language Model ScoreSunjae Kwon, Zonghai Yao, Harmon S. Jordan, David A. Levy 等EMNLP 2022 · 被引用 12 次
- ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability AssessmentTarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra 等EMNLP 2024 · 被引用 7 次
- Automated Lay Language Summarization of Biomedical Scientific ReviewsYue Guo, Wei Qiu, Yizhong Wang, Trevor CohenAAAI 2021 · 被引用 100 次
- MDEval: Evaluating and Enhancing Markdown Awareness in Large Language ModelsZhongpu Chen, Yinfeng Liu, Long Shi, Zhi-Jie Wang 等WWW 2025 · 被引用 12 次
- AutoMedEval: Harnessing Language Models for Automatic Medical Capability EvaluationXiechi Zhang, Zetian Ouyang, Linlin Wang, Gerard de Melo 等ACL 2025 · 被引用 1 次
