Lune

IEEE VR2026Top-tier venue

Multimodal Analysis of Speech-Gaze Fusion in Mixed Reality for the Detection of Neurodegenerative Disorders

Milosz Dudek, Jakub Sikora, Daria Hemmerling, Mateusz Daniol, Marek Wodzinski, Magdalena Wójcik-Pedziwiatr

2026Year

Abstract

Mixed reality (MR) headsets can synchronously capture eye movements and speech during ecological tasks, enabling interpretable, multimodal behavioural assessment. This study introduces an MR-native pipeline that fuses gaze and speech to characterize Parkinson’s disease (PD) in 10 PD patients and 18 healthy controls (HC) during a 40 s picture description on Microsoft HoloLens 2.Audio is transcribed and force-aligned with word-level timestamps, linguistically annotated and converted into lexical (tokens, unique tokens, MTLD (measure of textual lexical diversity)), syntactic (proper nouns per 100 tokens), fluency (words-per-minute), and pause-based temporal features. Gaze is filtered and summarized into kinematic measures (mean gaze speed, mean gaze acceleration, and acceleration variability) and fixation rate. Aligned gaze speech segments were independently rated for correspondence, yielding a per-participant alignment accuracy used in downstream analysis.Group contrasts use Mann–Whitney U, Cliff’s δ, and FDR control (global and family-wise) and show a distributional shift toward lower alignment in PD. Speech-derived markers (total/voiced words per minute, tokens, unique tokens, MTLD) are reduced in PD, gaze fixation rate also trends lower; proper nouns per 100 tokens is higher in PD, indicating a higher rate of proper-noun usage relative to transcript length in this task. A compact Top-K set (K=7) yields meaningful multivariate separability (centroid distance 2.826, 95% CI [1.900,4.113]) and nearest-centroid balanced accuracy 0.733, which further improves when adding alignment as an 8th feature (distance 2.854, CI [1.981,4.149]; accuracy 0.783).MR offers clear advantages over conventional setups: the headset co-registers gaze and speech in situ without external rigs, preserves ecological validity, and supports repeatable, low-burden, time-synchronized capture in clinics and at home. These findings indicate that MR gaze–speech fusion can capture complementary PD deficits and suggests a scalable path toward interpretable digital biomarkers. However, the conclusions are constrained by the limited sample size, and future validation in larger, independent cohorts is required to confirm generalizability and clinical utility.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 212712c4-7810-4bfe-9b92-3b04468ed5c6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines