Scalable Attentive Sentence Pair Modeling via Distilled Sentence Embedding
Oren Barkan, Noam Razin, Itzik Malkiel, Ori Katz, Avi Caciularu, Noam Koenigstein
Abstract
Recent state-of-the-art natural language understanding models, such as BERT and XLNet, score a pair of sentences (𝐴 and 𝐵) using multiple cross-attention operationsa process in which each word in sentence 𝐴 attends to all words in sentence 𝐵 and vice versa. As a result, computing the similarity between a query sentence and a set of candidate sentences, requires the propagation of all query-candidate sentencepairs throughout a stack of cross-attention layers. This exhaustive process becomes computationally prohibitive when the number of candidate sentences is large. In contrast, sentence embedding techniques learn a sentence-to-vector mapping and compute the similarity between the sentence vectors via simple elementary operations. In this paper, we introduce Distilled Sentence Embedding (DSE)a model that is based on knowledge distillation from cross-attentive models, focusing on sentence-pair tasks. The outline of DSE is as follows: Given a cross-attentive teacher model (e.g. a fine-tuned BERT), we train a sentence embedding based student model to reconstruct the sentence-pair scores obtained by the teacher model. We empirically demonstrate the effectiveness of DSE on five GLUE sentence-pair tasks. DSE significantly outperforms several ELMO variants and other sentence embedding methods, while accelerating computation of the query-candidate sentence-pairs similarities by several orders of magnitude, with an average relative degradation of 4.6% compared to BERT. Furthermore, we show that DSE produces sentence embeddings that reach state-of-the-art performance on universal sentence representation benchmarks. Our code is made publicly available at https://github.com/mi- crosoft/Distilled-Sentence-Embedding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef1420a7-f1a1-4e37-b211-eef414700b20Cited by top-tier papers3
- Interpreting BERT-based Text Similarity via Activation and Saliency MapsItzik Malkiel, Dvir Ginzburg, Oren Barkan, Avi Caciularu et al.WWW 2022 · 28 citations
- Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and BeyondOren Barkan, Tal Reiss, Jonathan Weill, Ori Katz et al.ICCV 2023 · 7 citations
- BEE: Metric-Adapted Explanations via Baseline Exploration-ExploitationOren Barkan, Yehonatan Elisha, Jonathan Weill, Noam KoenigsteinAAAI 2025
Related papers
- Static Word Embeddings for Sentence Semantic RepresentationTakashi Wada, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima et al.EMNLP 2025 · 1 citation
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep et al.EMNLP 2025
- Distilling Linguistic Context for Language Model CompressionGeondo Park, Gyeongman Kim, Eunho YangEMNLP 2021 · 24 citations
- Self-Guided Contrastive Learning for BERT Sentence RepresentationsTaeuk Kim, Kang Min Yoo, Sang-goo LeeACL 2021
