Learning Co-Speech Gesture for Multimodal Aphasia Type Detection
Daeun Lee, Sejung Son, Hyolim Jeon, Seungbae Kim, Jinyoung Han
摘要
Aphasia, a language disorder resulting from brain damage, requires accurate identification of specific aphasia types, such as Broca's and Wernicke's aphasia, for effective treatment. However, little attention has been paid to developing methods to detect different types of aphasia. Recognizing the importance of analyzing co-speech gestures for distinguish aphasia types, we propose a multimodal graph neural network for aphasia type detection using speech and corresponding gesture patterns. By learning the correlation between the speech and gesture modalities for each aphasia type, our model can generate textual representations sensitive to gesture information, leading to accurate aphasia type detection. Extensive experiments demonstrate the superiority of our approach over existing methods, achieving stateof-the-art results (F1 84.2%). We also show that gesture features outperform acoustic features, highlighting the significance of gesture expression in detecting aphasia types. We provide the codes for reproducibility purposes 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh 等ACL 2020 · 被引用 584 次
- Multi-Label Patent Categorization with Non-Local Attention-Based Graph Convolutional NetworkPingjie Tang, Meng Jiang, Bryan (Ning) Xia, Jed W. Pitera 等AAAI 2020 · 被引用 52 次
相关 Paper
- Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture GenerationFengqi Liu, Hexiang Wang, Jingyu Gong, Ran Yi 等ACM MM 2024 · 被引用 2 次
- Improving Handshape Representations for Sign Language Processing: A Graph Neural Network ApproachAlessa Carbo, Eric T. NalisnickEMNLP 2025
- MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture GenerationXiaofeng Mao, Zhengkai Jiang, Qilin Wang, Chencan Fu 等ACM MM 2024 · 被引用 7 次
- Understanding Co-Speech Gestures in-the-WildSindhu B. Hegde, K. R. Prajwal, Taein Kwon, Andrew ZissermanICCV 2025 · 被引用 4 次
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsSicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li 等ACM MM 2023 · 被引用 17 次
