Investigating Word-Class Distributions in Word Vector Spaces
Ryohei Sasano, Anna Korhonen
Abstract
This paper presents an investigation on the distribution of word vectors belonging to a certain word class in a pre-trained word vector space. To this end, we made several assumptions about the distribution, modeled the distribution accordingly, and validated each assumption by comparing the goodness of each model. Specifically, we considered two types of word classes -the semantic class of direct objects of a verb and the semantic class in a thesaurus -and tried to build models that properly estimate how likely it is that a word in the vector space is a member of a given word class. Our results on selectional preference and WordNet datasets show that the centroid-based model will fail to achieve good enough performance, the geometry of the distribution and the existence of subgroups will have limited impact, and also the negative instances need to be considered for adequate modeling of the distribution. We further investigated the relationship between the scores calculated by each model and the degree of membership and found that discriminative learning-based models are best in finding the boundaries of a class, while models based on the offset between positive and negative instances perform best in determining the degree of membership.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e790a0cd-51c7-484d-a41b-990dec5db666Cited by top-tier papers1
Ask how each one uses itRelated papers
- Multidirectional Associative Optimization of Function-Specific Word RepresentationsDaniela Gerz, Ivan Vulic, Marek Rei, Roi Reichart et al.ACL 2020
- Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary LearningXiaomeng Fan, Yuchuan Mao, Zhi Gao, Yuwei Wu et al.NeurIPS 2025 · 1 citation
- A Relation-Oriented Clustering Method for Open Relation ExtractionJun Zhao, Tao Gui, Qi Zhang, Yaqian ZhouEMNLP 2021 · 25 citations
- CATE: A Contrastive Pre-trained Model for Metaphor Detection with Semi-supervised LearningZhenxi Lin, Qianli Ma, Jiangyue Yan, Jieyu ChenEMNLP 2021 · 15 citations
- Rotation Has Two Sides: Evaluating Data Augmentation for Deep One-class ClassificationGuodong Wang, Yunhong Wang, Xiuguo Bao, Di HuangICLR 2024 · 3 citations
