Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
Fangyu Liu, Yunlong Jiao, Jordan Massiah, Emine Yilmaz, Serhii Havrylov
Abstract
In NLP, a large volume of tasks involve pairwise comparison between two sequences (e.g., sentence similarity and paraphrase identification). Predominantly, two formulations are used for sentence-pair tasks: bi-encoders and cross-encoders. Bi-encoders produce fixed-dimensional sentence representations and are computationally efficient, however, they usually underperform cross-encoders. Crossencoders can leverage their attention heads to exploit inter-sentence interactions for better performance but they require task finetuning and are computationally more expensive. In this paper, we present a completely unsupervised sentence-pair model termed as TRANS-ENCODER that combines the two learning paradigms into an iterative joint framework to simultaneously learn enhanced bi-and crossencoders. Specifically, on top of a pre-trained language model (PLM), we start with converting it to an unsupervised bi-encoder, and then alternate between the bi-and cross-encoder task formulations. In each alternation, one task formulation will produce pseudo-labels which are used as learning signals for the other task formulation. We then propose an extension to conduct such self-distillation approach on multiple PLMs in parallel and use the average of their pseudo-labels for mutual-distillation. TRANS-ENCODER creates, to the best of our knowledge, the first completely unsupervised cross-encoder and also a state-of-the-art unsupervised bi-encoder for sentence similarity. Both the bi-encoder and cross-encoder formulations of TRANS-ENCODER outperform recently proposed state-of-the-art unsupervised sentence encoders such as Mirror-BERT (Liu et al., 2021) and SimCSE (Gao et al., 2021) by up to 5% on the sentence similarity benchmarks. Code and models are released at https://github.com/amzn/trans-encoder .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72ddccd3-3153-456b-9787-8f4a96abed86Cited by top-tier papers9
- NN Prompting: Beyond-Context Learning with Calibration-Free Nearest Neighbor InferenceBenfeng Xu, Quan Wang, Zhendong Mao, Yajuan Lyu et al.ICLR 2023 · 12 citations
- Ranking-Enhanced Unsupervised Sentence Representation LearningYeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary et al.ACL 2023 · 12 citations
- Beyond Two-Tower Matching: Learning Sparse Retrievable Cross-Interactions for RecommendationLiangcai Su, Fan Yan, Jieming Zhu, Xi Xiao et al.SIGIR 2023 · 11 citations
- miCSE: Mutual Information Contrastive Learning for Low-shot Sentence EmbeddingsTassilo Klein, Moin NabiACL 2023 · 10 citations
- LEA: Improving Sentence Similarity Robustness to Typos Using Lexical Attention BiasMario Almagro, Emilio J. Almazán, Diego Ortego, David JiménezKDD 2023 · 6 citations
Builds on11
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 316 citations
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
- Semantic Re-tuning with Contrastive TensionFredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist et al.ICLR 2021 · 86 citations
Related papers
- Bilingual alignment transfers to multilingual alignment for unsupervised parallel text miningChih-chan Tien, Shane Steinert-ThrelkeldACL 2022 · 10 citations
- Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersFangyu Liu, Ivan Vulic, Anna Korhonen, Nigel CollierEMNLP 2021 · 85 citations
- Alleviating Over-smoothing for Unsupervised Sentence RepresentationNuo Chen, Linjun Shou, Jian Pei, Ming Gong et al.ACL 2023 · 10 citations
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation TransferYuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang et al.ACL 2021
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim et al.EMNLP 2020 · 126 citations
