What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
Aditya Yadavalli, Tiago Pimentel, Tamar I. Regev, Ethan Gotlieb Wilcox, Alex Warstadt
摘要
Prosody -the melody of speech -conveys critical information often not captured by the words or text of a message. In this paper, we propose an information-theoretic approach to quantify how much information is expressed by prosody alone and not by text, and crucially, what that information is about. Our approach applies large speech and language models to estimate the mutual information between a particular dimension of an utterance's meaning (e.g., its emotion) and any of its communication channels (e.g., audio or text). We then use this approach to quantify how much information is conveyed by audio and text about sarcasm, emotion, and questionhood, using speech from television and podcasts. We find that for sarcasm and emotion the audio channel -and by implication the prosodic channel -transmits over an order of magnitude more information about these features than the text channel alone, at least when long-term context beyond the current sentence is unavailable. For questionhood, prosody provides comparatively less additional information. We conclude by outlining a program applying our approach to more dimensions of meaning, communication channels, and languages. 1 However, the information conveyed by prosody is so critical to successful communication that humans have developed numerous typographical conventions to recover much of that information loss, including commas, question marks, italics, and emojis (Chafe, 1988; Holtgraves and Robinson, 2020) .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod 等ACL 2020 · 被引用 21 次
相关 Paper
- Quantifying the redundancy between prosody and textLukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell 等EMNLP 2023 · 被引用 5 次
- The Prosody of EmojisGiulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry HaddowACL 2026
- The time scale of redundancy between prosody and linguistic contextTamar I. Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf 等ACL 2025 · 被引用 5 次
- Your tone speaks louder than your face! Modality Order Infused Multi-modal Sarcasm DetectionMohit Tomar, Abhisek Tiwari, Tulika Saha, Sriparna SahaACM MM 2023 · 被引用 15 次
- Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-TuningLivia Qian, Gabriel SkantzeACL 2026
