Scalable Attentive Sentence Pair Modeling via Distilled Sentence Embedding
Oren Barkan, Noam Razin, Itzik Malkiel, Ori Katz, Avi Caciularu, Noam Koenigstein
摘要
Recent state-of-the-art natural language understanding models, such as BERT and XLNet, score a pair of sentences (𝐴 and 𝐵) using multiple cross-attention operationsa process in which each word in sentence 𝐴 attends to all words in sentence 𝐵 and vice versa. As a result, computing the similarity between a query sentence and a set of candidate sentences, requires the propagation of all query-candidate sentencepairs throughout a stack of cross-attention layers. This exhaustive process becomes computationally prohibitive when the number of candidate sentences is large. In contrast, sentence embedding techniques learn a sentence-to-vector mapping and compute the similarity between the sentence vectors via simple elementary operations. In this paper, we introduce Distilled Sentence Embedding (DSE)a model that is based on knowledge distillation from cross-attentive models, focusing on sentence-pair tasks. The outline of DSE is as follows: Given a cross-attentive teacher model (e.g. a fine-tuned BERT), we train a sentence embedding based student model to reconstruct the sentence-pair scores obtained by the teacher model. We empirically demonstrate the effectiveness of DSE on five GLUE sentence-pair tasks. DSE significantly outperforms several ELMO variants and other sentence embedding methods, while accelerating computation of the query-candidate sentence-pairs similarities by several orders of magnitude, with an average relative degradation of 4.6% compared to BERT. Furthermore, we show that DSE produces sentence embeddings that reach state-of-the-art performance on universal sentence representation benchmarks. Our code is made publicly available at https://github.com/mi- crosoft/Distilled-Sentence-Embedding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Interpreting BERT-based Text Similarity via Activation and Saliency MapsItzik Malkiel, Dvir Ginzburg, Oren Barkan, Avi Caciularu 等WWW 2022 · 被引用 28 次
- Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and BeyondOren Barkan, Tal Reiss, Jonathan Weill, Ori Katz 等ICCV 2023 · 被引用 7 次
- BEE: Metric-Adapted Explanations via Baseline Exploration-ExploitationOren Barkan, Yehonatan Elisha, Jonathan Weill, Noam KoenigsteinAAAI 2025
相关 Paper
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima 等EMNLP 2025 · 被引用 1 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep 等EMNLP 2025
- Distilling Linguistic Context for Language Model CompressionGeondo Park, Gyeongman Kim, Eunho YangEMNLP 2021 · 被引用 24 次
- Self-Guided Contrastive Learning for BERT Sentence RepresentationsTaeuk Kim, Kang Min Yoo, Sang-goo LeeACL 2021
