Revisiting the Optimality of Word Lengths
Tiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald, Ryan Cotterell
摘要
Zipf (1935) posited that wordforms are optimized to minimize utterances' communicative costs. Under the assumption that cost is given by an utterance's length, he supported this claim by showing that words' lengths are inversely correlated with their frequencies. Communicative cost, however, can be operationalized in different ways. Piantadosi et al. (2011) claim that cost should be measured as the distance between an utterance's information rate and channel capacity, which we dub the channel capacity hypothesis (CCH) here. Following this logic, they then proposed that a word's length should be proportional to the expected value of its surprisal (negative log-probability in context). In this work, we show that Piantadosi et al.'s derivation does not minimize CCH's cost, but rather a lower bound, which we term CCH ↓ . We propose a novel derivation, suggesting an improved way to minimize CCH's cost. Under this method, we find that a language's word lengths should instead be proportional to the surprisal's expectation plus its variance-tomean ratio. Experimentally, we compare these three communicative cost functions: Zipf's, CCH ↓ , and CCH. Across 13 languages and several experimental settings, we find that length is better predicted by frequency than either of the other hypotheses. In fact, when surprisal's expectation, or expectation plus variance-to-mean ratio, is estimated using better language models, it leads to worse word length predictions. We take these results as evidence that Zipf's longstanding hypothesis holds. https://github.com/tpimentelms/ optimality-of-word-lengths
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Quantifying the redundancy between prosody and textLukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell 等EMNLP 2023 · 被引用 5 次
- Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-AccentEthan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel 等ACL 2025 · 被引用 3 次
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 被引用 2 次
- More frequent verbs are associated with more diverse valency frames: Efficient principles at the lexicon-grammar interfaceSiyu Tao, Lucia Donatelli, Michael HahnACL 2024
它引用的顶会 Paper2
相关 Paper
- Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in DialogueThomas P. Utting, Mario Giulianelli, Arabella SinclairACL 2026
- Speakers enhance contextually confusable wordsEric Meinhardt, Eric Bakovic, Leon BergenACL 2020 · 被引用 36 次
- Lewis's Signaling Game as beta-VAE For Natural Word Lengths and SegmentsRyo Ueda, Tadahiro TaniguchiICLR 2024 · 被引用 13 次
- Surprisal Predicts Code-Switching in Chinese-English Bilingual TextJesús Calvillo, Le Fang, Jeremy R. Cole, David ReitterEMNLP 2020 · 被引用 6 次
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan 等ACL 2025 · 被引用 3 次
