Revisiting the Optimality of Word Lengths
Tiago Pimentel, Clara Meister, Ethan Wilcox, Kyle Mahowald, Ryan Cotterell
Abstract
Zipf (1935) posited that wordforms are optimized to minimize utterances' communicative costs. Under the assumption that cost is given by an utterance's length, he supported this claim by showing that words' lengths are inversely correlated with their frequencies. Communicative cost, however, can be operationalized in different ways. Piantadosi et al. (2011) claim that cost should be measured as the distance between an utterance's information rate and channel capacity, which we dub the channel capacity hypothesis (CCH) here. Following this logic, they then proposed that a word's length should be proportional to the expected value of its surprisal (negative log-probability in context). In this work, we show that Piantadosi et al.'s derivation does not minimize CCH's cost, but rather a lower bound, which we term CCH ↓ . We propose a novel derivation, suggesting an improved way to minimize CCH's cost. Under this method, we find that a language's word lengths should instead be proportional to the surprisal's expectation plus its variance-tomean ratio. Experimentally, we compare these three communicative cost functions: Zipf's, CCH ↓ , and CCH. Across 13 languages and several experimental settings, we find that length is better predicted by frequency than either of the other hypotheses. In fact, when surprisal's expectation, or expectation plus variance-to-mean ratio, is estimated using better language models, it leads to worse word length predictions. We take these results as evidence that Zipf's longstanding hypothesis holds. https://github.com/tpimentelms/ optimality-of-word-lengths
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bbd1b51-cddb-48ed-9e74-69e49c1b845eCited by top-tier papers4
- Quantifying the redundancy between prosody and textLukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell et al.EMNLP 2023 · 5 citations
- Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-AccentEthan Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel et al.ACL 2025 · 3 citations
- How to Compute the Probability of a WordTiago Pimentel, Clara MeisterEMNLP 2024 · 2 citations
- More frequent verbs are associated with more diverse valency frames: Efficient principles at the lexicon-grammar interfaceSiyu Tao, Lucia Donatelli, Michael HahnACL 2024
Builds on2
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
- A surprisal-duration trade-off across and within the world's languagesTiago Pimentel, Clara Meister, Elizabeth Salesky, Simone Teufel et al.EMNLP 2021 · 2 citations
Related papers
- Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in DialogueThomas P. Utting, Mario Giulianelli, Arabella SinclairACL 2026
- Speakers enhance contextually confusable wordsEric Meinhardt, Eric Bakovic, Leon BergenACL 2020 · 36 citations
- Lewis's Signaling Game as beta-VAE For Natural Word Lengths and SegmentsRyo Ueda, Tadahiro TaniguchiICLR 2024 · 13 citations
- Surprisal Predicts Code-Switching in Chinese-English Bilingual TextJesús Calvillo, Le Fang, Jeremy R. Cole, David ReitterEMNLP 2020 · 6 citations
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan et al.ACL 2025 · 3 citations
