RAW-C: Relatedness of Ambiguous Words in Context (A New Lexical Resource for English)
Sean Trott, Benjamin K. Bergen
摘要
Most words are ambiguous-they convey distinct meanings in different contexts-and even the meanings of unambiguous words are context-dependent. Both phenomena present a challenge for NLP. Recently, the advent of contextualized word embeddings has led to success on tasks involving lexical ambiguity, such as Word Sense Disambiguation. However, there are few tasks that directly evaluate how well these embeddings accommodate the continuous, dynamic nature of word meaningparticularly in a way that matches human intuitions. We introduce RAW-C, a dataset of graded, human relatedness judgments for 112 ambiguous words in context (with 672 sentence pairs total), as well as human estimates of sense dominance. The average inter-annotator agreement for the relatedness norms (assessed using a leave-one-annotatorout method) was 0.79. We then show that a measure of cosine distance, computed using contextualized embeddings from BERT and ELMo, correlates with human judgments, but that cosine distance also systematically underestimates how similar humans find uses of the same sense of a word to be, and systematically overestimates how similar humans find uses of different-sense homonyms. Finally, we propose a synthesis between psycholinguistic theories of the mental lexicon and computational models of lexical semantics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Context and POS in Action: A Comparative Study of Chinese Homonym Disambiguation in Human and Language ModelsChenwei Xie, Matthew King-Hang Ma, Wenbo Wang, William Shi-Yuan WangEMNLP 2025
- Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with PseudowordsTaelin Karidi, Yichu Zhou, Nathan Schneider, Omri Abend 等EMNLP 2021
它引用的顶会 Paper5
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 被引用 914 次
- Experience Grounds LanguageYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas 等EMNLP 2020 · 被引用 74 次
- CSI: A Coarse Sense Inventory for 85% Word Sense DisambiguationCaterina Lacerra, Michele Bevilacqua, Tommaso Pasini, Roberto NavigliAAAI 2020 · 被引用 31 次
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram 等ACL 2020 · 被引用 27 次
- (Re)construing Meaning in NLPSean Trott, Tiago Timponi Torrent, Nancy Chang, Nathan SchneiderACL 2020 · 被引用 2 次
相关 Paper
- Speakers Fill Lexical Semantic Gaps with ContextTiago Pimentel, Rowan Hall Maudslay, Damián E. Blasi, Ryan CotterellEMNLP 2020 · 被引用 1 次
- FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense DisambiguationMohamad Ballout, Anne Dedert, Nohayr Abdelmoneim, Ulf Krumnack 等EMNLP 2024 · 被引用 2 次
- Measuring Context-Word Biases in Lexical Semantic DatasetsQianchu Liu, Diana McCarthy, Anna KorhonenEMNLP 2022
- Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and SynonymyMarcos GarcíaACL 2021
- Can Embeddings Adequately Represent Medical Terminology? New Large-Scale Medical Term Similarity Datasets Have the Answer!Claudia Schulz, Damir JuricAAAI 2020 · 被引用 10 次
