A Corpus for Large-Scale Phonetic Typology
Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner, Ryan Cotterell, Alan W. Black, Jason Eisner
Abstract
A major hurdle in data-driven research on typology is having sufficient data in many languages to draw meaningful conclusions. We present VoxClamantis V1.0, the first largescale corpus for phonetic typology, with aligned segments and estimated phonemelevel labels in 690 readings spanning 635 languages, along with acoustic-phonetic measures of vowels and sibilants. Access to such data can greatly facilitate investigation of phonetic typology at a large scale and across many languages. However, it is nontrivial and computationally intensive to obtain such alignments for hundreds of languages, many of which have few to no resources presently available. We describe the methodology to create our corpus, discuss caveats with current methods and their impact on the utility of this data, and illustrate possible research directions through a series of case studies on the 48 highest-quality readings. Our corpus and scripts are publicly available for non-commercial use at https:// voxclamantisproject.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ZIPA: A family of efficient models for multilingual phone recognitionJian Zhu, Farhan Samir, Eleanor Chodroff, David R. MortensenACL 2025 · 10 citations
- A surprisal-duration trade-off across and within the world's languagesTiago Pimentel, Clara Meister, Elizabeth Salesky, Simone Teufel et al.EMNLP 2021 · 2 citations
Related papers
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu et al.ACL 2021
- Phonotomizer: A Compact, Unsupervised, Online Training Approach to Real-Time, Multilingual Phonetic SegmentationMichael S. Yantosca, Albert M. K. ChengACL 2025
- ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard TranslationKhoa Anh Ta, Nguyen Van Dinh, Kiet Van NguyenAAAI 2026
- CaMEL: Case Marker Extraction without LabelsLeonie Weissweiler, Valentin Hofmann, Masoud Jalili Sabet, Hinrich SchützeACL 2022 · 3 citations
- English-based acoustic models perform well in the forced alignment of two English-based Pacific CreolesSam Passmore, Lila San Roque, Kirsty Gillespie, Saurabh Nath et al.ACL 2025 · 1 citation
