RAW-C: Relatedness of Ambiguous Words in Context (A New Lexical Resource for English)
Sean Trott, Benjamin K. Bergen
Abstract
Most words are ambiguous-they convey distinct meanings in different contexts-and even the meanings of unambiguous words are context-dependent. Both phenomena present a challenge for NLP. Recently, the advent of contextualized word embeddings has led to success on tasks involving lexical ambiguity, such as Word Sense Disambiguation. However, there are few tasks that directly evaluate how well these embeddings accommodate the continuous, dynamic nature of word meaningparticularly in a way that matches human intuitions. We introduce RAW-C, a dataset of graded, human relatedness judgments for 112 ambiguous words in context (with 672 sentence pairs total), as well as human estimates of sense dominance. The average inter-annotator agreement for the relatedness norms (assessed using a leave-one-annotatorout method) was 0.79. We then show that a measure of cosine distance, computed using contextualized embeddings from BERT and ELMo, correlates with human judgments, but that cosine distance also systematically underestimates how similar humans find uses of the same sense of a word to be, and systematically overestimates how similar humans find uses of different-sense homonyms. Finally, we propose a synthesis between psycholinguistic theories of the mental lexicon and computational models of lexical semantics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd3896a4-9892-4758-985d-491141127be8Cited by top-tier papers2
- Context and POS in Action: A Comparative Study of Chinese Homonym Disambiguation in Human and Language ModelsChenwei Xie, Matthew King-Hang Ma, Wenbo Wang, William Shi-Yuan WangEMNLP 2025
- Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with PseudowordsTaelin Karidi, Yichu Zhou, Nathan Schneider, Omri Abend et al.EMNLP 2021
Builds on5
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Experience Grounds LanguageYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas et al.EMNLP 2020 · 74 citations
- CSI: A Coarse Sense Inventory for 85% Word Sense DisambiguationCaterina Lacerra, Michele Bevilacqua, Tommaso Pasini, Roberto NavigliAAAI 2020 · 31 citations
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram et al.ACL 2020 · 27 citations
- (Re)construing Meaning in NLPSean Trott, Tiago Timponi Torrent, Nancy Chang, Nathan SchneiderACL 2020 · 2 citations
Related papers
- Speakers Fill Lexical Semantic Gaps with ContextTiago Pimentel, Rowan Hall Maudslay, Damián E. Blasi, Ryan CotterellEMNLP 2020 · 1 citation
- FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense DisambiguationMohamad Ballout, Anne Dedert, Nohayr Abdelmoneim, Ulf Krumnack et al.EMNLP 2024 · 2 citations
- Measuring Context-Word Biases in Lexical Semantic DatasetsQianchu Liu, Diana McCarthy, Anna KorhonenEMNLP 2022
- Exploring the Representation of Word Meanings in Context: A Case Study on Homonymy and SynonymyMarcos GarcíaACL 2021
- Can Embeddings Adequately Represent Medical Terminology? New Large-Scale Medical Term Similarity Datasets Have the Answer!Claudia Schulz, Damir JuricAAAI 2020 · 10 citations
