Sentence Smith: Controllable Edits for Evaluating Text Embeddings
Hongji Li, Andrianos Michail, Reto Gubelmann, Simon Clematide, Juri Opitz
摘要
Controllable and transparent text generation has been a long-standing goal in NLP. Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbolic representation, and generating from it. However, earlier approaches were hindered by parsing and generation insufficiencies. Using modern parsers and a safety supervision mechanism, we show how close current methods come to this goal. Concretely, we propose the Sentence Smith framework for English, which has three steps: 1. Parsing a sentence into a semantic graph. 2. Applying human-designed semantic manipulation rules. 3. Generating text from the manipulated graph. A final entailment check (4.) verifies the validity of the applied transformation. To demonstrate our framework's utility, we use it to induce hard negative text pairs that challenge text embedding models. Since the controllable generation makes it possible to clearly isolate different types of semantic shifts, we can evaluate text embedding models in a fine-grained way, also addressing an issue in current benchmarking where linguistic phenomena remain opaque. Human validation confirms that our transparent generation process produces texts of good quality. Notably, our way of generation is very resource-efficient, since it relies only on smaller neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 被引用 137 次
相关 Paper
- T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted AttackBoxin Wang, Hengzhi Pei, Boyuan Pan, Qian Chen 等EMNLP 2020 · 被引用 55 次
- Disentangled Learning with Synthetic Parallel Data for Text Style TransferJingxuan Han, Quan Wang, Zikang Guo, Benfeng Xu 等ACL 2024 · 被引用 4 次
- Bridging Continuous and Discrete Spaces: Interpretable Sentence Representation Learning via Compositional OperationsJames Y. Huang, Wenlin Yao, Kaiqiang Song, Hongming Zhang 等EMNLP 2023
- Sentence Representation Learning with Generative Objective rather than Contrastive ObjectiveBohong Wu, Hai ZhaoEMNLP 2022 · 被引用 3 次
- HGM³: Hierarchical Generative Masked Motion Modeling with Hard Token MiningMinjae Jeong, Yechan Hwang, Jaejin Lee, Sungyoon Jung 等ICLR 2025
