Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment
Ryo Nagata, Hiroya Takamura, Naoki Otani, Yoshifumi Kawasaki
摘要
In this paper, we propose methods 1 for discovering semantic differences in words appearing in two corpora. The key idea is to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector, which is equivalent to examining a kind of variance of the word vector distribution. The proposed methods do not require alignments between words and/or corpora for comparison that previous methods do. All they require are to compute variance (or norms of mean word vectors) for each word type. Nevertheless, they rival the best-performing system in the SemEval-2020 Task 1. In addition, they are (i) robust for the skew in corpus sizes; (ii) capable of detecting semantic differences in infrequent words; and (iii) effective in pinpointing word instances that have a meaning missing in one of the two corpora under comparison. We show these advantages for historical corpora and also for native/non-native English corpora.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Quantifying Lexical Semantic Shift via Unbalanced Optimal TransportRyo Kishino, Hiroaki Yamagiwa, Ryo Nagata, Sho Yokoi 等ACL 2025
- Verifiable LLM-Generated Text Detection via Projected Semantic-Structural DistributionsRuochong Xiong, Qien Li, Wangwang Lian, Yulong Wan 等ACL 2026
- A New Formulation of Zipf's Meaning-Frequency Law through Contextual DiversityRyo Nagata, Kumiko Tanaka-IshiiACL 2025
它引用的顶会 Paper2
- Analysing Lexical Semantic Change with Contextualised Word RepresentationsMario Giulianelli, Marco Del Tredici, Raquel FernándezACL 2020 · 被引用 118 次
- Simple, Interpretable and Stable Method for Detecting Words with Usage Change across CorporaHila Gonen, Ganesh Jawahar, Djamé Seddah, Yoav GoldbergACL 2020 · 被引用 57 次
相关 Paper
- Accurate and Efficient Statistical Testing for Word Semantic BreadthYo EharaACL 2026
- Fake it Till You Make it: Self-Supervised Semantic Shifts for Monolingual Word Embedding TasksMaurício Gruppi, Pin-Yu Chen, Sibel AdaliAAAI 2021 · 被引用 7 次
- Current Semantic-change Quantification Methods Struggle with Discovery in the WildKhonzoda Umarova, Lillian Lee, Laerdon KimEMNLP 2025
- Norm of Word Embedding Encodes Information GainMomose Oyama, Sho Yokoi, Hidetoshi ShimodairaEMNLP 2023 · 被引用 6 次
- Word Rotator's DistanceSho Yokoi, Ryo Takahashi, Reina Akama, Jun Suzuki 等EMNLP 2020
