Re-evaluating Word Mover's Distance
Ryoma Sato, Makoto Yamada, Hisashi Kashima
摘要
The word mover's distance (WMD) is a fundamental technique for measuring the similarity of two documents. As the crux of WMD, it can take advantage of the underlying geometry of the word space by employing an optimal transport formulation. The original study on WMD reported that WMD outperforms classical baselines such as bag-of-words (BOW) and TF-IDF by significant margins in various datasets. In this paper, we point out that the evaluation in the original study could be misleading. We re-evaluate the performances of WMD and the classical baselines and find that the classical baselines are competitive with WMD if we employ an appropriate preprocessing, i.e., L1 normalization. In addition, we introduce an analogy between WMD and L1-normalized BOW and find that not only the performance of WMD but also the distance values resemble those of BOW in high dimensional spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon 等ICLR 2023 · 被引用 8 次
- Characterizing and Measuring Linguistic Dataset DriftTyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas 等ACL 2023 · 被引用 2 次
- A linear time approximation of Wasserstein distance with word embedding selectionSho Otao, Makoto YamadaEMNLP 2023 · 被引用 2 次
- Zero-Shot Task Adaptation with Relevant Feature InformationAtsutoshi Kumagai, Tomoharu Iwata, Yasuhiro FujiwaraAAAI 2024 · 被引用 1 次
- DiNaM: Disinformation Narrative Mining with Large Language ModelsWitold Sosnowski, Arkadiusz Modzelewski, Kinga Skorupska, Adam WierzbickiEMNLP 2025
它引用的顶会 Paper15
- A Fair Comparison of Graph Neural Networks for Graph ClassificationFederico Errica, Marco Podda, Davide Bacciu, Alessio MicheliICLR 2020 · 被引用 508 次
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 被引用 167 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Projection Robust Wasserstein Distance and Riemannian OptimizationTianyi Lin, Chenyou Fan, Nhat Ho, Marco Cuturi 等NeurIPS 2020 · 被引用 84 次
- Message Passing Attention Networks for Document UnderstandingGiannis Nikolentzos, Antoine J.-P. Tixier, Michalis VazirgiannisAAAI 2020 · 被引用 80 次
相关 Paper
- Word Rotator's DistanceSho Yokoi, Ryo Takahashi, Reina Akama, Jun Suzuki 等EMNLP 2020
- Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity MetricIvan P. Yamshchikov, Viacheslav Shibaev, Nikolay Khlebnikov, Alexey TikhonovAAAI 2021 · 被引用 38 次
- OTLDA: A Geometry-aware Optimal Transport Approach for Topic ModelingViet Huynh, He Zhao, Dinh PhungNeurIPS 2020 · 被引用 29 次
- Scalable Sobolev IPM for Probability Measures on a GraphTam Le, Truyen Nguyen, Hideitsu Hino, Kenji FukumizuICML 2025
- Wasserstein Distance Regularized Sequence Representation for Text Matching in Asymmetrical DomainsWeijie Yu, Chen Xu, Jun Xu, Liang Pang 等EMNLP 2020 · 被引用 10 次
