ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers
Yi Liu, Yuekang Li, Gelei Deng, Felix Juefei-Xu, Yao Du, Cen Zhang, Chengwei Liu, Yeting Li, Lei Ma, Yang Liu
摘要
The popularity of automatic speech recognition (ASR) systems nowadays leads to an increasing need for improving their accessibility. Handling stuttering speech is an important feature for accessible ASR systems. To improve the accessibility of ASR systems for stutterers, we need to expose and analyze the failures of ASR systems on stuttering speech. The speech datasets recorded from stutterers are not diverse enough to expose most of the failures. Furthermore, these datasets lack ground truth information about the non-stuttered text, rendering them unsuitable as comprehensive test suites. Therefore, a methodology for generating stuttering speech as test inputs to test and analyze the performance of ASR systems is needed. However, generating valid test inputs in this scenario is challenging. The reason is that although the generated test inputs should mimic how stutterers speak, they should also be diverse enough to trigger more failures. To address the challenge, we propose Aster, a technique for automatically testing the accessibility of ASR systems. Aster can generate valid test cases by injecting five different types of stuttering. The generated test cases can both simulate realistic stuttering speech and expose failures in ASR systems. Moreover, Aster can further enhance the quality of the test cases with a multi-objective optimization-based seed updating algorithm. We implemented Aster as a framework and evaluated it on four open-source ASR models and three commercial ASR systems. We conduct a comprehensive evaluation of Aster and find that it significantly increases the word error rate, match error rate, and word information loss in the evaluated ASR systems. Additionally, our user study demonstrates that the generated stuttering audio is indistinguishable from real-world stuttering audio clips.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Structure-invariant testing for machine translationPinjia He, Clara Meister, Zhendong SuICSE 2020 · 被引用 84 次
- Latte: Use-Case and Assistive-Service Driven Automated Accessibility Testing Framework for AndroidNavid Salehnamadi, Abdulaziz Alshayban, Jun-Wei Lin, Iftekhar Ahmed 等CHI 2021 · 被引用 52 次
- Semantic Web Accessibility Testing via Hierarchical Visual AnalysisMohammad Bajammal, Ali MesbahICSE 2021 · 被引用 29 次
相关 Paper
- Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who StutterXinru Tang, Jingjin Li, Shaomei WuCHI 2026 · 被引用 3 次
- From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech RecognitionColin Lea, Zifang Huang, Jaya Narain, Lauren Tooley 等CHI 2023 · 被引用 35 次
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin 等NeurIPS 2024 · 被引用 14 次
- ASRTest: automated testing for deep-neural-network-driven speech recognition systemsPin Ji, Yang Feng, Jia Liu, Zhihong Zhao 等ISSTA 2022 · 被引用 22 次
- Interventional Speech Noise Injection for ASR Generalizable Spoken Language UnderstandingYeonJoon Jung, Jaeseong Lee, Seungtaek Choi, Dohyeon Lee 等EMNLP 2024
