Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring
Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, Jason Weston
Abstract
The use of deep pre-trained transformers has led to remarkable progress in a number of applications (Devlin et al., 2018). For tasks that make pairwise comparisons between sequences, matching a given input with a corresponding label, two approaches are common: Cross-encoders performing full self-attention over the pair and Bi-encoders encoding the pair separately. The former often performs better, but is too slow for practical use. In this work, we develop a new transformer architecture, the Poly-encoder, that learns global rather than token level self-attention features. We perform a detailed comparison of all three approaches, including what pre-training and fine-tuning strategies work best. We show our models achieve state-of-the-art results on four tasks; that Poly-encoders are faster than Cross-encoders and more accurate than Bi-encoders; and that the best results are obtained by pre-training on large datasets similar to the downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6affebdc-0551-4199-8fbd-8188ee895ab2Cited by top-tier papers57
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Hash Layers For Large Sparse ModelsStephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, Jason WestonNeurIPS 2021 · 316 citations
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 200 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Towards Persona-Based Empathetic Conversational ModelsPeixiang Zhong, Chen Zhang, Hao Wang, Yong Liu et al.EMNLP 2020 · 112 citations
Related papers
- Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillationsFangyu Liu, Yunlong Jiao, Jordan Massiah, Emine Yilmaz et al.ICLR 2022 · 36 citations
- Multi-Level Head-Wise Match and Aggregation in Transformer for Textual Sequence MatchingShuohang Wang, Yunshi Lan, Yi Tay, Jing Jiang et al.AAAI 2020 · 8 citations
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek et al.EMNLP 2020 · 268 citations
- MEXMA: Token-level objectives improve sentence representationsJoão Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari, Loïc BarraultACL 2025
- Cross-Thought for Sentence Encoder Pre-trainingShuohang Wang, Yuwei Fang, Siqi Sun, Zhe Gan et al.EMNLP 2020 · 17 citations
