Is the Same Performance Really the Same?: Understanding How Listeners Perceive ASR Results Differently According to the Speaker's Accent
Seoyoung Kim, Yeon Su Park, Dakyeom Ahn, Jin Myung Kwak, Juho Kim
Abstract
Research suggests that automatic speech recognition (ASR) systems, which automatically convert speech to text, show different performances according to various input classes (e.g., accent, age), requiring attention to building fairer AI systems that would perform similarly across various input classes. However, would an AI system with the same performance regardless of input classes really be perceived as fair enough? To this end, we investigate how listeners perceive the ASR system of the same result differently according to whether the speaker is a native speaker (NS) or a non-native speaker (NNS), which may lead to unfair situations. We conducted a study (n = 420), where participants were given one of the ten speech recordings with various accents of the same script along with the same captions. We found that even with the same ASR output, listeners perceive the ASR results differently. They found captions to be more useful for NNS's speech and blamed NNS more for the errors than NS. Based on the findings, we present design implications suggesting that we should take a step further than just achieving the same performance across various input classes to build a fair ASR system.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7e1a240e-d182-4f61-84a6-98ea9e9ac1c9Cited by top-tier papers1
Ask how each one uses itRelated papers
- "It feels like we're not meeting the criteria": Examining and Mitigating the Cascading Effects of Bias in Automatic Speech Recognition in Spoken Language InterfacesKelechi Ezema, Chelsea Chandler, Rosy Southwell, Niranjan Cholendiran et al.CHI 2025 · 8 citations
- Classist Tools: Social Class Correlates with Performance in NLPAmanda Cercas Curry, Giuseppe Attanasio, Zeerak Talat, Dirk HovyACL 2024 · 1 citation
- AI-Based Speaking Assistant: Supporting Non-Native Speakers' Speaking in Real-Time Multilingual CommunicationPeinuan Qin, Zicheng Zhu, Naomi Yamashita, Yitian Yang et al.CSCW 2025 · 3 citations
- Can Voice Assistants Be Microaggressors? Cross-Race Psychological Responses to Failures of Automatic Speech RecognitionKimi Wenzel, Nitya Devireddy, Cam Davidson, Geoff KaufmanCHI 2023 · 24 citations
- Using Phoneme Representations to Build Predictive Models Robust to ASR ErrorsAnjie Fang, Simone Filice, Nut Limsopatham, Oleg RokhlenkoSIGIR 2020 · 16 citations
