Quantifying the redundancy between prosody and text
Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, Tamar Regev
摘要
Prosody-the suprasegmental component of speech, including pitch, loudness, and tempocarries critical aspects of meaning. However, the relationship between the information conveyed by prosody vs. by the words themselves remains poorly understood. We use large language models (LLMs) to estimate how much information is redundant between prosody and the words themselves. Using a large spoken corpus of English audiobooks, we extract prosodic features aligned to individual words and test how well they can be predicted from LLM embeddings, compared to non-contextual word embeddings. We find a high degree of redundancy between the information carried by the words and prosodic information across several prosodic features, including intensity, duration, pauses, and pitch contours. Furthermore, a word's prosodic information is redundant with both the word itself and the context preceding as well as following it. Still, we observe that prosodic features can not be fully predicted from text, suggesting that prosody carries information above and beyond the words. Along with this paper, we release a general-purpose data processing pipeline for quantifying the relationship between linguistic information and extra-linguistic features. https://github.com/lu-wo/ quantifying-redundancy
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The time scale of redundancy between prosody and linguistic contextTamar I. Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf 等ACL 2025 · 被引用 5 次
- Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-AccentEthan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel 等ACL 2025 · 被引用 3 次
- What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple ChannelsAditya Yadavalli, Tiago Pimentel, Tamar I. Regev, Ethan Gotlieb Wilcox 等ACL 2026 · 被引用 1 次
- The Prosody of EmojisGiulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry HaddowACL 2026
- Modeling Bottom-up Information Quality during Language ProcessingCui Ding, Yanning Yin, Lena Ann Jäger, Ethan WilcoxEMNLP 2025
它引用的顶会 Paper5
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod 等ACL 2020 · 被引用 21 次
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang 等ACL 2023 · 被引用 13 次
- Revisiting the Optimality of Word LengthsTiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald 等EMNLP 2023 · 被引用 5 次
- A surprisal-duration trade-off across and within the world's languagesTiago Pimentel, Clara Meister, Elizabeth Salesky, Simone Teufel 等EMNLP 2021 · 被引用 2 次
相关 Paper
- A Variational Framework for Improving Naturalness in Generative Spoken Language ModelsLi-Wei Chen, Takuya Higuchi, Zakaria Aldeneh, Ahmed Hussen Abdelaziz 等ICML 2025
- Learning de-identified representations of prosody from raw audioJack Weston, Raphael Lenain, Udeepa Meepegama, Emil FristedICML 2021 · 被引用 18 次
- VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic AugmentationYancheng Wang, Osama Hanna, Ruiming Xie, Xianfeng Rui 等ICLR 2026 · 被引用 4 次
- Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMsDingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang 等EMNLP 2025 · 被引用 1 次
- Prosodic segmentation for parsing spoken dialogueElizabeth Nielsen, Mark Steedman, Sharon GoldwaterACL 2021
