Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
Fangyu Liu, Yunlong Jiao, Jordan Massiah, Emine Yilmaz, Serhii Havrylov
摘要
In NLP, a large volume of tasks involve pairwise comparison between two sequences (e.g., sentence similarity and paraphrase identification). Predominantly, two formulations are used for sentence-pair tasks: bi-encoders and cross-encoders. Bi-encoders produce fixed-dimensional sentence representations and are computationally efficient, however, they usually underperform cross-encoders. Crossencoders can leverage their attention heads to exploit inter-sentence interactions for better performance but they require task finetuning and are computationally more expensive. In this paper, we present a completely unsupervised sentence-pair model termed as TRANS-ENCODER that combines the two learning paradigms into an iterative joint framework to simultaneously learn enhanced bi-and crossencoders. Specifically, on top of a pre-trained language model (PLM), we start with converting it to an unsupervised bi-encoder, and then alternate between the bi-and cross-encoder task formulations. In each alternation, one task formulation will produce pseudo-labels which are used as learning signals for the other task formulation. We then propose an extension to conduct such self-distillation approach on multiple PLMs in parallel and use the average of their pseudo-labels for mutual-distillation. TRANS-ENCODER creates, to the best of our knowledge, the first completely unsupervised cross-encoder and also a state-of-the-art unsupervised bi-encoder for sentence similarity. Both the bi-encoder and cross-encoder formulations of TRANS-ENCODER outperform recently proposed state-of-the-art unsupervised sentence encoders such as Mirror-BERT (Liu et al., 2021) and SimCSE (Gao et al., 2021) by up to 5% on the sentence similarity benchmarks. Code and models are released at https://github.com/amzn/trans-encoder .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- NN Prompting: Beyond-Context Learning with Calibration-Free Nearest Neighbor InferenceBenfeng Xu, Quan Wang, Zhendong Mao, Yajuan Lyu 等ICLR 2023 · 被引用 12 次
- Ranking-Enhanced Unsupervised Sentence Representation LearningYeon Seonwoo, Guoyin Wang, Changmin Seo, Sajal Choudhary 等ACL 2023 · 被引用 12 次
- Beyond Two-Tower Matching: Learning Sparse Retrievable Cross-Interactions for RecommendationLiangcai Su, Fan Yan, Jieming Zhu, Xi Xiao 等SIGIR 2023 · 被引用 11 次
- miCSE: Mutual Information Contrastive Learning for Low-shot Sentence EmbeddingsTassilo Klein, Moin NabiACL 2023 · 被引用 10 次
- LEA: Improving Sentence Similarity Robustness to Typos Using Lexical Attention BiasMario Almagro, Emilio J. Almazán, Diego Ortego, David JiménezKDD 2023 · 被引用 6 次
它引用的顶会 Paper11
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence ScoringSamuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason WestonICLR 2020 · 被引用 316 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- Semantic Re-tuning with Contrastive TensionFredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist 等ICLR 2021 · 被引用 86 次
相关 Paper
- Bilingual alignment transfers to multilingual alignment for unsupervised parallel text miningChih-chan Tien, Shane Steinert-ThrelkeldACL 2022 · 被引用 10 次
- Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence EncodersFangyu Liu, Ivan Vulic, Anna Korhonen, Nigel CollierEMNLP 2021 · 被引用 85 次
- Alleviating Over-smoothing for Unsupervised Sentence RepresentationNuo Chen, Linjun Shou, Jian Pei, Ming Gong 等ACL 2023 · 被引用 10 次
- ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation TransferYuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang 等ACL 2021
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim 等EMNLP 2020 · 被引用 126 次
