Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora
Hila Gonen, Ganesh Jawahar, Djamé Seddah, Yoav Goldberg
Abstract
The problem of comparing two bodies of text and searching for words that differ in their usage between them arises often in digital humanities and computational social science. This is commonly approached by training word embeddings on each corpus, aligning the vector spaces, and looking for words whose cosine distance in the aligned space is large. However, these methods often require extensive filtering of the vocabulary to perform well, and-as we show in this work-result in unstable, and hence less reliable, results. We propose an alternative approach that does not use vector space alignment, and instead considers the neighbors of each word. The method is simple, interpretable and stable. We demonstrate its effectiveness in 9 different setups, considering different corpus splitting criteria (age, gender and profession of tweet authors, time of tweet) and different languages (English, French and Hebrew).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97fc7e08-e5ab-4ca2-96c9-2045d60cdd0cCited by top-tier papers5
- Discovering Differences in the Representation of People using Contextualized Semantic AxesLi Lucy, Divya Tadimeti, David BammanEMNLP 2022 · 7 citations
- Narrative Characteristics in Refugee Discourse: An Analysis of American Public Opinion on the Afghan Refugee Crisis After the Taliban TakeoverHulya Dogan, Kiet A. Nguyen, Ismini LourentzouCSCW 2024 · 3 citations
- Variance Matters: Detecting Semantic Differences without Corpus/Word AlignmentRyo Nagata, Hiroya Takamura, Naoki Otani, Yoshifumi KawasakiEMNLP 2023 · 2 citations
- Time is Encoded in the Weights of Finetuned Language ModelsKai Nylund, Suchin Gururangan, Noah A. SmithACL 2024
- Modeling the Evolution of English Noun Compounds with Feature-Rich Diachronic Compositionality PredictionFilip Miletic, Sabine Schulte im WaldeACL 2025
Related papers
- Measure and Evaluation of Semantic Divergence across Two LanguagesSyrielle Montariol, Alexandre AllauzenACL 2021
- Interpretable Text Embeddings and Text Similarity Explanation: A SurveyJuri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó et al.EMNLP 2025 · 3 citations
- Fake it Till You Make it: Self-Supervised Semantic Shifts for Monolingual Word Embedding TasksMaurício Gruppi, Pin-Yu Chen, Sibel AdaliAAAI 2021 · 7 citations
- Word Rotator's DistanceSho Yokoi, Ryo Takahashi, Reina Akama, Jun Suzuki et al.EMNLP 2020
- Analyzing the Surprising Variability in Word Embedding Stability Across LanguagesLaura Burdick, Jonathan K. Kummerfeld, Rada MihalceaEMNLP 2021 · 8 citations
