A Corpus for Large-Scale Phonetic Typology
Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner, Ryan Cotterell, Alan W. Black, Jason Eisner
摘要
A major hurdle in data-driven research on typology is having sufficient data in many languages to draw meaningful conclusions. We present VoxClamantis V1.0, the first largescale corpus for phonetic typology, with aligned segments and estimated phonemelevel labels in 690 readings spanning 635 languages, along with acoustic-phonetic measures of vowels and sibilants. Access to such data can greatly facilitate investigation of phonetic typology at a large scale and across many languages. However, it is nontrivial and computationally intensive to obtain such alignments for hundreds of languages, many of which have few to no resources presently available. We describe the methodology to create our corpus, discuss caveats with current methods and their impact on the utility of this data, and illustrate possible research directions through a series of case studies on the 48 highest-quality readings. Our corpus and scripts are publicly available for non-commercial use at https:// voxclamantisproject.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ZIPA: A family of efficient models for multilingual phone recognitionJian Zhu, Farhan Samir, Eleanor Chodroff, David R. MortensenACL 2025 · 被引用 10 次
- A surprisal-duration trade-off across and within the world's languagesTiago Pimentel, Clara Meister, Elizabeth Salesky, Simone Teufel 等EMNLP 2021 · 被引用 2 次
相关 Paper
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu 等ACL 2021
- Phonotomizer: A Compact, Unsupervised, Online Training Approach to Real-Time, Multilingual Phonetic SegmentationMichael S. Yantosca, Albert M. K. ChengACL 2025
- ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard TranslationKhoa Anh Ta, Nguyen Van Dinh, Kiet Van NguyenAAAI 2026
- CaMEL: Case Marker Extraction without LabelsLeonie Weissweiler, Valentin Hofmann, Masoud Jalili Sabet, Hinrich SchützeACL 2022 · 被引用 3 次
- English-based acoustic models perform well in the forced alignment of two English-based Pacific CreolesSam Passmore, Lila San Roque, Kirsty Gillespie, Saurabh Nath 等ACL 2025 · 被引用 1 次
