The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
Fagun Patel, Duc Q. Nguyen, Sang T. Truong, Jody Vaynshtok, Sanmi Koyejo, Nick Haber
Abstract
According to the U.S. National Institutes of Health, more than 3.4 million children experience speech disorders that require clinical intervention. The number of speech-language pathologists (SLPs) is roughly 20 times fewer than the number of affected children, highlighting a significant gap in children's care and a pressing need for technological support that improves the productivity of SLPs. State-ofthe-art multimodal language models (MLMs) show promise for supporting SLPs, but their use remains underexplored largely due to a limited understanding of their performance in highstakes clinical settings. To address this gap, we collaborate with domain experts to develop a taxonomy of real-world use cases of MLMs in speech-language pathologies. Building on this taxonomy, we introduce the first comprehensive benchmark for evaluating MLM across five core use cases, each containing 1,000 manually annotated data points. This benchmark includes robustness and sensitivity tests under various settings, including background noise, speaker gender, and accent. Our evaluation of 15 state-of-the-art MLMs reveals that no single model consistently outperforms others across all tasks. Notably, we find systematic disparities, with models performing better on male speakers, and observe that chain-of-thought prompting can degrade performance on classification tasks with large label spaces and narrow decision boundaries. Furthermore, we study fine-tuning MLMs on domain-specific data, achieving improvements of over 10% compared to base models. These findings highlight both the potential and limitations of current MLMs for speech-language pathology applications, underscoring the need for further research and targeted development 1 . 1 To support continued progress, we publicly release our datasets, fine-tuned models, and benchmarking framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0408ca0e-5d0c-4af6-99a1-de8df7238fb7Builds on3
- Iterative Reasoning Preference OptimizationRichard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho et al.NeurIPS 2024 · 287 citations
- Continual Pre-training of Language ModelsZixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi et al.ICLR 2023 · 15 citations
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans WorseRyan Liu, Jiayi Geng, Addison J. Wu, Ilia Sucholutsky et al.ICML 2025 · 4 citations
Related papers
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang et al.ICLR 2026 · 143 citations
- Exploring AI-Based Support in Speech-Language Pathology for Culturally and Linguistically Diverse ChildrenAaleyah Lewis, Aayushi Dangol, Hyewon Suh, Abbie Olszewski et al.CHI 2025 · 19 citations
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?Yifan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu et al.ICLR 2025 · 1 citation
- U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound UnderstandingAnjie Le, Henan Liu, Yue Wang, Zhenyu Liu et al.ICLR 2026 · 8 citations
- Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsJie Liu, Wenxuan Wang, Yihang Su, Jingyuan Huang et al.ACL 2025 · 16 citations
