Linguistically Conditioned Semantic Textual Similarity
Jingxuan Tu, Keer Xu, Liulu Yue, Bingyang Ye, Kyeongmin Rim, James Pustejovsky
摘要
Semantic textual similarity (STS) is a fundamental NLP task that measures the semantic similarity between a pair of sentences. In order to reduce the inherent ambiguity posed from the sentences, a recent work called Conditional STS (C-STS) has been proposed to measure the sentences' similarity conditioned on a certain aspect. Despite the popularity of C-STS, we find that the current C-STS dataset suffers from various issues that could impede proper evaluation on this task. In this paper, we reannotate the C-STS validation set and observe an annotator discrepancy on 55% of the instances resulting from the annotation errors in the original label, ill-defined conditions, and the lack of clarity in the task definition. After a thorough dataset analysis, we improve the C-STS task by leveraging the models' capability to understand the conditions under a QA task setting. With the generated answers, we present an automatic error identification pipeline that is able to identify annotation errors from the C-STS data with over 80% F1 score. We also propose a new method that largely improves the performance over baselines on the C-STS data by training the models with the answers. Finally we discuss the conditionality annotation based on the typed-feature structure (TFS) of entity types. We show in examples that the TFS is able to provide a linguistic foundation for constructing C-STS data with new conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Enhanced Noun-Noun Compound Interpretation through Textual EnrichmentBingyang Ye, Jingxuan Tu, James PustejovskyEMNLP 2025
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
它引用的顶会 Paper4
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Why LLMs Hallucinate, and How to Get (Evidential) Closure: Perceptual, Intensional, and Extensional Learning for Faithful Natural Language GenerationAdam BouyamournEMNLP 2023 · 被引用 9 次
- C-STS: Conditional Semantic Textual SimilarityAmeet Deshpande, Carlos E. Jimenez, Howard Chen, Vishvak Murahari 等EMNLP 2023 · 被引用 9 次
- Leveraging QA Datasets to Improve Generative Data AugmentationDheeraj Mekala, Tu Vu, Timo Schick, Jingbo ShangEMNLP 2022 · 被引用 8 次
相关 Paper
- PoLi-RL: A Point-to-List Reinforcement Learning Framework for Conditional Semantic Textual SimilarityZixin Song, Bowen Zhang, Qian-Wen Zhang, Di Yin 等ICLR 2026 · 被引用 1 次
- Identifying Ambiguous Similarity Conditions via Semantic MatchingHan-Jia Ye, Yi Shi, De-Chuan ZhanCVPR 2022 · 被引用 5 次
- CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question AnsweringZongxi Li, Yang Li, Haoran Xie, S. Joe QinEMNLP 2025
- TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language ModelsJie He, Bo Peng, Yi Liao, Qun Liu 等ACL 2021
- ConditionalQA: A Complex Reading Comprehension Dataset with Conditional AnswersHaitian Sun, William W. Cohen, Ruslan SalakhutdinovACL 2022 · 被引用 41 次
