ZIPA: A family of efficient models for multilingual phone recognition
Jian Zhu, Farhan Samir, Eleanor Chodroff, David R. Mortensen
摘要
We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPAPack++, a large-scale multilingual speech corpus with 17,132 hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation. With the large-scale training data, ZIPA, including transducer (ZIPA-T) and CTC-based (ZIPA-CR) variants, leverage the efficient Zipformer backbones and outperform existing phone recognition systems with much fewer parameters. Further scaling via noisy student training on 11,000 hours of pseudo-labeled multilingual data yields further improvement. While ZIPA achieves strong performance on benchmarks, error analysis reveals persistent limitations in modeling sociophonetic diversity, underscoring challenges for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- POWSM: A Phonetic Open Whisper-Style Speech Foundation ModelChin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo 等ACL 2026 · 被引用 11 次
- PRiSM: Benchmarking Phone Realization in Speech ModelsShikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim, Kwanghee Choi 等ACL 2026 · 被引用 3 次
- Towards Language-Agnostic STIPA: Universal Phonetic Transcription to Support Language Documentation at ScaleJacob Lee Suchardt, Hana El-Shazli, Pierluigi CassottiEMNLP 2025
- LingGym: How Far Are LLMs from Thinking Like Field Linguists?Changbing Yang, Franklin Ma, Freda Shi, Jian ZhuEMNLP 2025
它引用的顶会 Paper10
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and UnderstandingYifan Peng, Siddharth Dalmia, Ian R. Lane, Shinji WatanabeICML 2022 · 被引用 203 次
- Zipformer: A faster and better encoder for automatic speech recognitionZengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang 等ICLR 2024 · 被引用 155 次
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li 等EMNLP 2024 · 被引用 19 次
相关 Paper
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech RecognitionTianduo Wang, Lu Xu, Wei Lu, Shanbo ChengEMNLP 2025 · 被引用 1 次
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani 等ICML 2021 · 被引用 140 次
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu 等ACL 2021
- Phonotomizer: A Compact, Unsupervised, Online Training Approach to Real-Time, Multilingual Phonetic SegmentationMichael S. Yantosca, Albert M. K. ChengACL 2025
- Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic AlphabetMilan Miletic, Julie Kallini, Ekaterina ShutovaACL 2026
