GPT-D: Inducing Dementia-related Linguistic Anomalies by Deliberate Degradation of Artificial Neural Language Models
Changye Li, David S. Knopman, Weizhe Xu, Trevor Cohen, Serguei Pakhomov
Abstract
Deep learning (DL) techniques involving fine-tuning large numbers of model parameters have delivered impressive performance on the task of discriminating between language produced by cognitively healthy individuals, and those with Alzheimer’s disease (AD). However, questions remain about their ability to generalize beyond the small reference sets that are publicly available for research. As an alternative to fitting model parameters directly, we propose a novel method by which a Transformer DL model (GPT-2) pre-trained on general English text is paired with an artificially degraded version of itself (GPT-D), to compute the ratio between these two models’ perplexities on language from cognitively healthy and impaired individuals. This technique approaches state-of-the-art performance on text data from a widely used “Cookie Theft” picture description task, and unlike established alternatives also generalizes well to spontaneous conversations. Furthermore, GPT-D generates text with characteristics known to be associated with AD, demonstrating the induction of dementia-related linguistic anomalies. Our study is a step toward better understanding of the relationships between the inner workings of generative neural language models, the language that they produce, and the deleterious effects of dementia on human speech and language characteristics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Towards Domain-Agnostic and Domain-Adaptive Dementia Detection from Spoken LanguageShahla Farzana, Natalie PardeACL 2023 · 5 citations
- Mitigating Confounding in Speech-Based Dementia Detection through Weight MaskingZhecheng Sheng, Xiruo Ding, Brian Hur, Changye Li et al.ACL 2025 · 1 citation
Builds on3
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Roles and Utilization of Attention Heads in Transformer-based Neural Language ModelsJae-young Jo, Sung-Hyon MyaengACL 2020 · 32 citations
- A Tale of Two Perplexities: Sensitivity of Neural Language Models to Lexical Retrieval Deficits in Dementia of the Alzheimer's TypeTrevor Cohen, Serguei PakhomovACL 2020 · 2 citations
Related papers
- SPZ: A Semantic Perturbation-based Data Augmentation Method with Zonal-Mixing for Alzheimer's Disease DetectionFangfang Li, Cheng Huang, Puzhen Su, Jie YinACL 2024 · 1 citation
- DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's DiseaseTingyu Mo, Jacqueline C. K. Lam, Victor O. K. Li, Lawrence Y. L. CheungAAAI 2025 · 4 citations
- Neural Language Models are not Born Equal to Fit Brain Data, but Training HelpsAlexandre Pasquiou, Yair Lakretz, John T. Hale, Bertrand Thirion et al.ICML 2022 · 44 citations
- Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with DementiaDimitris Gkoumas, Matthew Purver, Maria LiakataEMNLP 2023 · 4 citations
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
