Multimodal Analysis of Speech-Gaze Fusion in Mixed Reality for the Detection of Neurodegenerative Disorders
Milosz Dudek, Jakub Sikora, Daria Hemmerling, Mateusz Daniol, Marek Wodzinski, Magdalena Wójcik-Pedziwiatr
摘要
Mixed reality (MR) headsets can synchronously capture eye movements and speech during ecological tasks, enabling interpretable, multimodal behavioural assessment. This study introduces an MR-native pipeline that fuses gaze and speech to characterize Parkinson’s disease (PD) in 10 PD patients and 18 healthy controls (HC) during a 40 s picture description on Microsoft HoloLens 2.Audio is transcribed and force-aligned with word-level timestamps, linguistically annotated and converted into lexical (tokens, unique tokens, MTLD (measure of textual lexical diversity)), syntactic (proper nouns per 100 tokens), fluency (words-per-minute), and pause-based temporal features. Gaze is filtered and summarized into kinematic measures (mean gaze speed, mean gaze acceleration, and acceleration variability) and fixation rate. Aligned gaze speech segments were independently rated for correspondence, yielding a per-participant alignment accuracy used in downstream analysis.Group contrasts use Mann–Whitney U, Cliff’s δ, and FDR control (global and family-wise) and show a distributional shift toward lower alignment in PD. Speech-derived markers (total/voiced words per minute, tokens, unique tokens, MTLD) are reduced in PD, gaze fixation rate also trends lower; proper nouns per 100 tokens is higher in PD, indicating a higher rate of proper-noun usage relative to transcript length in this task. A compact Top-K set (K=7) yields meaningful multivariate separability (centroid distance 2.826, 95% CI [1.900,4.113]) and nearest-centroid balanced accuracy 0.733, which further improves when adding alignment as an 8th feature (distance 2.854, CI [1.981,4.149]; accuracy 0.783).MR offers clear advantages over conventional setups: the headset co-registers gaze and speech in situ without external rigs, preserves ecological validity, and supports repeatable, low-burden, time-synchronized capture in clinics and at home. These findings indicate that MR gaze–speech fusion can capture complementary PD deficits and suggests a scalable path toward interpretable digital biomarkers. However, the conclusions are constrained by the limited sample size, and future validation in larger, independent cohorts is required to confirm generalizability and clinical utility.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SkiMR: Dwell-free Eye Typing in Mixed RealityJinghui Hu, John J. Dudley, Per Ola KristenssonIEEE VR 2024 · 被引用 6 次
- Understood: Real-Time Communication Support for Adults with ADHD Using Mixed RealityShizhen Zhang, Shengxin Li, Quan LiUIST 2025 · 被引用 2 次
- Using Speech to Visualise Shared Gaze Cues in MR Remote CollaborationAllison Jing, Gun A. Lee, Mark BillinghurstIEEE VR 2022 · 被引用 21 次
- Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with DementiaDimitris Gkoumas, Matthew Purver, Maria LiakataEMNLP 2023 · 被引用 4 次
- Contactless Upper-Limb Bradykinesia Monitoring for Parkinson's Disease via Semantic-Aware mmWave Sensing in Daily LifeJinjian Wang, Qingyong Hu, Yizhen Zhang, Yuxuan Zhou 等UbiComp 2026
