Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
Ethan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel, Alex Warstadt, Tamar I. Regev
Abstract
This paper argues that the relationship between lexical identity and prosody -- one well-studied parameter of linguistic variation -- can be characterized using information theory. We predict that languages that use prosody to make lexical distinctions should exhibit a higher mutual information between word identity and prosody, compared to languages that don't. We test this hypothesis in the domain of pitch, which is used to make lexical distinctions in tonal languages, like Cantonese. We use a dataset of speakers reading sentences aloud in ten languages across five language families to estimate the mutual information between the text and their pitch curves. We find that, across languages, pitch curves display similar amounts of entropy. However, these curves are easier to predict given their associated text in the tonal languages, compared to pitch- and stress-accent languages, and thus the mutual information is higher in these languages, supporting our hypothesis. Our results support perspectives that view linguistic typology as gradient, rather than categorical.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07987ab3-bf39-4af7-9874-8839b98a6bbdCited by top-tier papers3
- The time scale of redundancy between prosody and linguistic contextTamar I. Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf et al.ACL 2025 · 5 citations
- What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple ChannelsAditya Yadavalli, Tiago Pimentel, Tamar I. Regev, Ethan Gotlieb Wilcox et al.ACL 2026 · 1 citation
- Modeling Bottom-up Information Quality during Language ProcessingCui Ding, Yanning Yin, Lena Ann Jäger, Ethan WilcoxEMNLP 2025
Builds on4
- Revisiting the Optimality of Word LengthsTiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald et al.EMNLP 2023 · 5 citations
- Quantifying the redundancy between prosody and textLukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell et al.EMNLP 2023 · 5 citations
- The time scale of redundancy between prosody and linguistic contextTamar I. Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf et al.ACL 2025 · 5 citations
- Causal Estimation of Tokenisation BiasPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos et al.ACL 2025
Related papers
- An Information-Theoretic Foundation for the Subregular HierarchyMai Phan Quoc Hung, Khanh Nguyen Quoc, Doàn Minh Luong, Duong Thu Ngan et al.ACL 2026
- Norm of Word Embedding Encodes Information GainMomose Oyama, Sho Yokoi, Hidetoshi ShimodairaEMNLP 2023 · 6 citations
- ToneCraft: Cantonese Lyrics Generation with Harmony of Tones and PitchesJunyu Cheng, Chang Pan, Shuangyin LiEMNLP 2025
- Speakers Fill Lexical Semantic Gaps with ContextTiago Pimentel, Rowan Hall Maudslay, Damián E. Blasi, Ryan CotterellEMNLP 2020 · 1 citation
- A Cognitive Regularizer for Language ModelingJason Wei, Clara Meister, Ryan CotterellACL 2021
