MedReadMe: A Systematic Study for Fine-grained Sentence Readability in Medical Domain
Chao Jiang, Wei Xu
Abstract
Medical texts are notoriously challenging to read. Properly measuring their readability is the first step towards making them more accessible. In this paper, we present a systematic study on fine-grained readability measurements in the medical domain at both sentence-level and span-level. We introduce a new dataset MedReadMe, which consists of manually annotated readability ratings and fine-grained complex span annotation for 4,520 sentences, featuring two novel "Google-Easy" and "Google-Hard" categories. It supports our quantitative analysis, which covers 650 linguistic features and automatic complex word and jargon identification. Enabled by our high-quality annotation, we benchmark and improve several state-of-the-art sentence-level readability metrics for the medical domain specifically, which include unsupervised, supervised, and prompting-based methods using recently developed large language models (LLMs). Informed by our fine-grained complex span annotation, we find that adding a single feature, capturing the number of jargon spans, into existing readability formulas can significantly improve their correlation with human judgments. We will publicly release the dataset and code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e70b026d-fdef-4002-b9ad-c16dca88ea55Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong et al.ACL 2020 · 103 citations
- Expertise Style Transfer: A New Task Towards Better Communication between Experts and LaymenYixin Cao, Ruihao Shui, Liangming Pan, Min-Yen Kan et al.ACL 2020 · 50 citations
- Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic FeaturesBruce W. Lee, Yoo Sung Jang, Jason Hyung-Jong LeeEMNLP 2021 · 46 citations
- Making Science Simple: Corpora for the Lay Summarisation of Scientific LiteratureTomas Goldsack, Zhihao Zhang, Chenghua Lin, Carolina ScartonEMNLP 2022 · 38 citations
- LENS: A Learnable Evaluation Metric for Text SimplificationMounica Maddela, Yao Dou, David Heineman, Wei XuACL 2023 · 21 citations
Related papers
- MedJEx: A Medical Jargon Extraction Model with Wiki's Hyperlink Span and Contextualized Masked Language Model ScoreSunjae Kwon, Zonghai Yao, Harmon S. Jordan, David A. Levy et al.EMNLP 2022 · 12 citations
- ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability AssessmentTarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra et al.EMNLP 2024 · 7 citations
- Automated Lay Language Summarization of Biomedical Scientific ReviewsYue Guo, Wei Qiu, Yizhong Wang, Trevor CohenAAAI 2021 · 100 citations
- MDEval: Evaluating and Enhancing Markdown Awareness in Large Language ModelsZhongpu Chen, Yinfeng Liu, Long Shi, Zhi-Jie Wang et al.WWW 2025 · 12 citations
- AutoMedEval: Harnessing Language Models for Automatic Medical Capability EvaluationXiechi Zhang, Zetian Ouyang, Linlin Wang, Gerard de Melo et al.ACL 2025 · 1 citation
