Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora
Hila Gonen, Ganesh Jawahar, Djamé Seddah, Yoav Goldberg
摘要
The problem of comparing two bodies of text and searching for words that differ in their usage between them arises often in digital humanities and computational social science. This is commonly approached by training word embeddings on each corpus, aligning the vector spaces, and looking for words whose cosine distance in the aligned space is large. However, these methods often require extensive filtering of the vocabulary to perform well, and-as we show in this work-result in unstable, and hence less reliable, results. We propose an alternative approach that does not use vector space alignment, and instead considers the neighbors of each word. The method is simple, interpretable and stable. We demonstrate its effectiveness in 9 different setups, considering different corpus splitting criteria (age, gender and profession of tweet authors, time of tweet) and different languages (English, French and Hebrew).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Discovering Differences in the Representation of People using Contextualized Semantic AxesLi Lucy, Divya Tadimeti, David BammanEMNLP 2022 · 被引用 7 次
- Narrative Characteristics in Refugee Discourse: An Analysis of American Public Opinion on the Afghan Refugee Crisis After the Taliban TakeoverHulya Dogan, Kiet A. Nguyen, Ismini LourentzouCSCW 2024 · 被引用 3 次
- Variance Matters: Detecting Semantic Differences without Corpus/Word AlignmentRyo Nagata, Hiroya Takamura, Naoki Otani, Yoshifumi KawasakiEMNLP 2023 · 被引用 2 次
- Time is Encoded in the Weights of Finetuned Language ModelsKai Nylund, Suchin Gururangan, Noah A. SmithACL 2024
- Modeling the Evolution of English Noun Compounds with Feature-Rich Diachronic Compositionality PredictionFilip Miletic, Sabine Schulte im WaldeACL 2025
相关 Paper
- Measure and Evaluation of Semantic Divergence across Two LanguagesSyrielle Montariol, Alexandre AllauzenACL 2021
- Interpretable Text Embeddings and Text Similarity Explanation: A SurveyJuri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó 等EMNLP 2025 · 被引用 3 次
- Fake it Till You Make it: Self-Supervised Semantic Shifts for Monolingual Word Embedding TasksMaurício Gruppi, Pin-Yu Chen, Sibel AdaliAAAI 2021 · 被引用 7 次
- Word Rotator's DistanceSho Yokoi, Ryo Takahashi, Reina Akama, Jun Suzuki 等EMNLP 2020
- Analyzing the Surprising Variability in Word Embedding Stability Across LanguagesLaura Burdick, Jonathan K. Kummerfeld, Rada MihalceaEMNLP 2021 · 被引用 8 次
