Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMs
Yu-Wen Chen, Melody Ma, Julia Hirschberg
摘要
Automatic pronunciation assessment is typically performed by acoustic models trained on audio-score pairs. Although effective, these systems provide only numerical scores, without the information needed to help learners understand their errors. Meanwhile, large language models (LLMs) have proven effective in supporting language learning, but their potential for assessing pronunciation remains unexplored. In this work, we introduce TextPA, a zero-shot, Textual description-based Pronunciation Assessment approach. TextPA utilizes human-readable representations of speech signals, which are fed into an LLM to assess pronunciation accuracy and fluency, while also providing reasoning behind the assigned scores. Finally, a phoneme sequence match scoring method is used to refine the accuracy scores. Our work highlights a previously overlooked direction for pronunciation assessment. Instead of relying on supervised training with audioscore examples, we exploit the rich pronunciation knowledge embedded in written text. Experimental results show that our approach is both cost-efficient and competitive in performance. Furthermore, TextPA significantly improves the performance of conventional audioscore-trained models on out-of-domain data by offering a complementary perspective.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and DescriptionsSiyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian 等ACL 2025 · 被引用 20 次
- Zero-shot Large Language Models for Automatic Readability AssessmentRiley Grossman, Yi ChenACL 2026 · 被引用 1 次
- An Effective Pronunciation Assessment Approach Leveraging Hierarchical Transformers and Pre-training StrategiesBi-Cheng Yan, Jiun-Ting Li, Yi-Cheng Wang, Hsin-Wei Wang 等ACL 2024
- Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR TranscriptsJiaqing Liu, Chong Deng, Qinglin Zhang, Shilin Zhou 等AAAI 2025 · 被引用 1 次
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis 等ICCV 2025 · 被引用 3 次
