Speakers Fill Lexical Semantic Gaps with Context
Tiago Pimentel, Rowan Hall Maudslay, Damián E. Blasi, Ryan Cotterell
摘要
Lexical ambiguity is widespread in language, allowing for the reuse of economical word forms and therefore making language more efficient. If ambiguous words cannot be disambiguated from context, however, this gain in efficiency might make language less clear-resulting in frequent miscommunication. For a language to be clear and efficiently encoded, we posit that the lexical ambiguity of a word type should correlate with how much information context provides about it, on average. To investigate whether this is the case, we operationalise the lexical ambiguity of a word as the entropy of meanings it can take, and provide two ways to estimate this-one which requires human annotation (using Word-Net), and one which does not (using BERT), making it readily applicable to a large number of languages. We validate these measures by showing that, on six high-resource languages, there are significant Pearson correlations between our BERT-based estimate of ambiguity and the number of synonyms a word has in WordNet (e.g. ρ = 0.40 in English). We then test our main hypothesis-that a word's lexical ambiguity should negatively correlate with its contextual uncertainty-and find significant correlations on all 18 typologically diverse languages we analyse. This suggests that, in the presence of ambiguity, speakers compensate by making contexts more informative.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Measuring Context-Word Biases in Lexical Semantic DatasetsQianchu Liu, Diana McCarthy, Anna KorhonenEMNLP 2022
- An Unsupervised, Geometric and Syntax-aware Quantification of PolysemyAnmol Goel, Charu Sharma, Ponnurangam KumaraguruEMNLP 2022
它引用的顶会 Paper1
相关 Paper
- RAW-C: Relatedness of Ambiguous Words in Context (A New Lexical Resource for English)Sean Trott, Benjamin K. BergenACL 2021
- Analysing Lexical Semantic Change with Contextualised Word RepresentationsMario Giulianelli, Marco Del Tredici, Raquel FernándezACL 2020 · 被引用 118 次
- Speakers enhance contextually confusable wordsEric Meinhardt, Eric Bakovic, Leon BergenACL 2020 · 被引用 36 次
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram 等ACL 2020 · 被引用 27 次
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 被引用 2 次
