Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity Metric
Ivan P. Yamshchikov, Viacheslav Shibaev, Nikolay Khlebnikov, Alexey Tikhonov
摘要
The rapid development of such natural language processing tasks as style transfer, paraphrase, and machine translation often calls for the use of semantic similarity metrics. In recent years a lot of methods to measure the semantic similarity of two short texts were developed. This paper provides a comprehensive analysis for more than a dozen of such methods. Using a new dataset of fourteen thousand sentence pairs human-labeled according to their semantic similarity, we demonstrate that none of the metrics widely used in the literature is close enough to human judgment in these tasks. A number of recently proposed metrics provide comparable results, yet Word Mover Distance is shown to be the most reasonable solution to measure semantic similarity in reformulated texts at the moment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language GenerationMingkai Deng, Bowen Tan, Zhengzhong Liu, Eric P. Xing 等EMNLP 2021 · 被引用 49 次
- Text Detoxification using Large Pre-trained Neural ModelsDavid Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva 等EMNLP 2021 · 被引用 16 次
- Characterizing and Measuring Linguistic Dataset DriftTyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas 等ACL 2023 · 被引用 2 次
- T-STAR: Truthful Style Transfer using AMR Graph as Intermediate RepresentationAnubhav Jangra, Preksha Nema, Aravindan RaghuveerEMNLP 2022 · 被引用 1 次
相关 Paper
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 被引用 9 次
- Re-evaluating Word Mover's DistanceRyoma Sato, Makoto Yamada, Hisashi KashimaICML 2022 · 被引用 25 次
- Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language ModelsMirac Suzgun, Luke Melas-Kyriazi, Dan JurafskyEMNLP 2022 · 被引用 34 次
- Just Rank: Rethinking Evaluation with Word and Sentence SimilaritiesBin Wang, C.-C. Jay Kuo, Haizhou LiACL 2022 · 被引用 33 次
- HUME: Measuring the Human-Model Performance Gap in Text Embedding TasksAdnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff 等ICLR 2026 · 被引用 8 次
