Re-evaluating Word Mover's Distance
Ryoma Sato, Makoto Yamada, Hisashi Kashima
Abstract
The word mover's distance (WMD) is a fundamental technique for measuring the similarity of two documents. As the crux of WMD, it can take advantage of the underlying geometry of the word space by employing an optimal transport formulation. The original study on WMD reported that WMD outperforms classical baselines such as bag-of-words (BOW) and TF-IDF by significant margins in various datasets. In this paper, we point out that the evaluation in the original study could be misleading. We re-evaluate the performances of WMD and the classical baselines and find that the classical baselines are competitive with WMD if we employ an appropriate preprocessing, i.e., L1 normalization. In addition, we introduce an analogy between WMD and L1-normalized BOW and find that not only the performance of WMD but also the distance values resemble those of BOW in high dimensional spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90dc18d4-6a0c-4729-ac07-3f5ccfb016b1Cited by top-tier papers5
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon et al.ICLR 2023 · 8 citations
- Characterizing and Measuring Linguistic Dataset DriftTyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas et al.ACL 2023 · 2 citations
- A linear time approximation of Wasserstein distance with word embedding selectionSho Otao, Makoto YamadaEMNLP 2023 · 2 citations
- Zero-Shot Task Adaptation with Relevant Feature InformationAtsutoshi Kumagai, Tomoharu Iwata, Yasuhiro FujiwaraAAAI 2024 · 1 citation
- DiNaM: Disinformation Narrative Mining with Large Language ModelsWitold Sosnowski, Arkadiusz Modzelewski, Kinga Skorupska, Adam WierzbickiEMNLP 2025
Builds on15
- A Fair Comparison of Graph Neural Networks for Graph ClassificationFederico Errica, Marco Podda, Davide Bacciu, Alessio MicheliICLR 2020 · 508 citations
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 167 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Projection Robust Wasserstein Distance and Riemannian OptimizationTianyi Lin, Chenyou Fan, Nhat Ho, Marco Cuturi et al.NeurIPS 2020 · 84 citations
- Message Passing Attention Networks for Document UnderstandingGiannis Nikolentzos, Antoine J.-P. Tixier, Michalis VazirgiannisAAAI 2020 · 80 citations
Related papers
- Word Rotator's DistanceSho Yokoi, Ryo Takahashi, Reina Akama, Jun Suzuki et al.EMNLP 2020
- Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity MetricIvan P. Yamshchikov, Viacheslav Shibaev, Nikolay Khlebnikov, Alexey TikhonovAAAI 2021 · 38 citations
- OTLDA: A Geometry-aware Optimal Transport Approach for Topic ModelingViet Huynh, He Zhao, Dinh PhungNeurIPS 2020 · 29 citations
- Scalable Sobolev IPM for Probability Measures on a GraphTam Le, Truyen Nguyen, Hideitsu Hino, Kenji FukumizuICML 2025
- Wasserstein Distance Regularized Sequence Representation for Text Matching in Asymmetrical DomainsWeijie Yu, Chen Xu, Jun Xu, Liang Pang et al.EMNLP 2020 · 10 citations
