Lune

CHI2026Top-tier venue

Challenges in Automatic Speech Recognition for Adults with Cognitive Impairment

Michelle Cohn, Alyssa Lanzi, Yui Ishihara, Chen-Nee Chuah, Georgia Zellou, Alyssa Weakley

2026Year
2Citations

Abstract

Millions of people live with cognitive impairment from Alzheimer's disease and related dementias (ADRD). Voice-enabled smart home systems offer promise for supporting daily living but rely on automatic speech recognition (ASR) to transcribe their speech to text. Prior work has shown reduced ASR performance for adults with cognitive impairment; however, the acoustic factors underlying these disparities remain poorly understood. This paper evaluates ASR performance for 83 older adults across cognitive groups (cognitively normal, mild cognitive impairment, dementia) reading commands to a voice assistant (Amazon Alexa). Results show that ASR errors are significantly higher for individuals with dementia, revealing a critical usability gap. To better understand these disparities, we conducted an acoustic analysis of speech features and found that a speaker's intensity, voice quality, and pause ratio predicted ASR accuracy. Based on these findings, we outline HCI design implications for AgeTech and voice interfaces, including speaker-personalized ASR, human-in-the-loop correction of ASR transcripts, and interaction-level personalization to support ability-based adaptation.

which suggests that fine-tuning ASR models with these features could lead to improved accessibility.

Keywords : automatic speech recognition (ASR), mild cognitive impairment, dementia, Alzheimer's disease (AD), accessible technology

Millions of people currently have Alzheimer's disease and related dementias (ADRD) and this number is increasing, with 13.8 million cases anticipated by 2060 in the United States [101] . There is a rapid development of systems to support aging-in-place for older adults with cognitive impairment (AgeTech), such as smart home systems, home helper robots, and voice assistants [26,47,60,66,81,84,92,100] . AgeTech systems can screen for mild cognitive impairment (MCI), a precursor to ADRD for many individuals, such as based on language [46, 90] and movement [51] . AgeTech can also support activities of daily living (ADL) [92] and social interaction [72] . Many of these systems use voice assistants, which can provide an accessible interface that does not require expertise with mobile and computer interfaces [37,67] .

In systems using spoken interaction (e.g., voice assistants, home health robots), the technology transcribes the user's speech to text via automatic speech recognition (ASR). While ASR systems continue to improve, they still often show worse transcription for speech produced by older adults [61,87,94] , speakers with non-"standard" accents and dialects [14,21,38] , and individuals with neurological symptoms that shape language (e.g., dysarthria [70] , Down Syndrome [15] ). Compared to cognitively normal controls, individuals with cognitive impairment produce acoustic differences in their speech: a slower rate, reduced intensity, lower and less variable pitch, hoarser voice quality, and more frequent pauses [3,5,28,54,58] , which could impact ASR.

While there is a large body of work classifying cognitive status from speech (for a review, see Saeedi et al., 2024 ), few report the error rates for ASR and those that do exclusively do so as a proposed feature in detecting MCI/dementia [44,46,48,93,98,99] . Accordingly, we know little about when and why ASR errors occur for adults with cognitive impairment.

Understanding these errors has important implications for Human Computer Interaction (HCI) frameworks. First, ASR errors meaningfully shape users' interactions with voice-based AgeTech systems, impacting trust [9] , satisfaction [59] , and adoption [45] . Early misrecognition can have lasting impacts: if a person experiences errors from a system early on, they are less likely to use it --even if it later improves [16,50] . Second, disparities in ASR based on cognitive status can speak to Ability-Based Design (ABD) [95] , which argues that technologies should adapt to users' abilities, rather than requiring them to compensate for system limitations. In the context of cognitive impairment, this includes accommodating known speech production changes associated with decline. An ability-based perspective underscores the need to understand when and why ASR disparities arise to design voice-based AgeTech that is accessible and equitable.

This paper provides the following contributions:

• Empirical evidence that a state-of-the-art ASR model (Whisper) shows lower transcription accuracy for older adults with dementia.

• Acoustic analysis demonstrating that speech characteristics associated with cognitive impairment (e.g., reduced intensity, increased shimmer) result in higher ASR error rates.

• Design implications for inclusive AgeTech for cognitive impairment, including personalizing ASR models, human-in-the-loop support from users and caregivers, and interaction-level adaptation informed by dementia-related speech differences (e.g., extended turn taking).

There is a growing body of research developing systems

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f99823bf-c265-410c-a729-e48ca97a8e25

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines