A Tale of Two Perplexities: Sensitivity of Neural Language Models to Lexical Retrieval Deficits in Dementia of the Alzheimer's Type
Trevor Cohen, Serguei Pakhomov
Abstract
In recent years there has been a burgeoning interest in the use of computational methods to distinguish between elicited speech samples produced by patients with dementia, and those from healthy controls. The difference between perplexity estimates from two neural language models (LMs) -one trained on transcripts of speech produced by healthy participants and the other trained on transcripts from patients with dementia -as a single feature for diagnostic classification of unseen transcripts has been shown to produce state-of-the-art performance. However, little is known about why this approach is effective, and on account of the lack of case/control matching in the most widely-used evaluation set of transcripts (De-mentiaBank), it is unclear if these approaches are truly diagnostic, or are sensitive to other variables. In this paper, we interrogate neural LMs trained on participants with and without dementia using synthetic narratives previously developed to simulate progressive semantic dementia by manipulating lexical frequency. We find that perplexity of neural LMs is strongly and differentially associated with lexical frequency, and that a mixture model resulting from interpolating control and dementia LMs improves upon the current state-of-the-art for models trained on transcript text exclusively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 697ceba0-13e3-46d0-9b6e-d9154669c574Cited by top-tier papers3
- GPT-D: Inducing Dementia-related Linguistic Anomalies by Deliberate Degradation of Artificial Neural Language ModelsChangye Li, David S. Knopman, Weizhe Xu, Trevor Cohen et al.ACL 2022 · 24 citations
- Adversarial Text Generation using Large Language Models for Dementia DetectionYouxiang Zhu, Nana Lin, Kiran Balivada, Daniel Haehn et al.EMNLP 2024 · 1 citation
- Mitigating Confounding in Speech-Based Dementia Detection through Weight MaskingZhecheng Sheng, Xiruo Ding, Brian Hur, Changye Li et al.ACL 2025 · 1 citation
Related papers
- DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's DiseaseTingyu Mo, Jacqueline C. K. Lam, Victor O. K. Li, Lawrence Y. L. CheungAAAI 2025 · 4 citations
- Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with DementiaDimitris Gkoumas, Matthew Purver, Maria LiakataEMNLP 2023 · 4 citations
- A Digital Language Coherence Marker for Monitoring DementiaDimitris Gkoumas, Adam Tsakalidis, Maria LiakataEMNLP 2023 · 1 citation
- Towards Domain-Agnostic and Domain-Adaptive Dementia Detection from Spoken LanguageShahla Farzana, Natalie PardeACL 2023 · 5 citations
- Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze SurprisalSathvik Nair, Byung-Doh OhACL 2026 · 1 citation
