Look Both Ways and No Sink: Converting LLMs into Text Encoders without Training
Ziyong Lin, Haoyi Wu, Shu Wang, Kewei Tu, Zilong Zheng, Zixia Jia
Abstract
Recent advancements have demonstrated the advantage of converting pretrained large language models into powerful text encoders by enabling bidirectional attention in transformer layers. However, existing methods often require extensive training on large-scale datasets, posing challenges in low-resource, domain-specific scenarios. In this work, we show that a pretrained large language model can be converted into a strong text encoder without additional training. We first conduct a comprehensive empirical study to investigate different conversion strategies and identify the impact of the attention sink phenomenon on the performance of converted encoder models. Based on our findings, we propose a novel approach that enables bidirectional attention and suppresses the attention sink phenomenon, resulting in superior performance. Extensive experiments on multiple domains demonstrate the effectiveness of our approach. Our work provides new insights into the training-free conversion of text encoders in low-resource scenarios and contributes to the advancement of domain-specific text representation generation. Our code is available at https://github.com/bigai-nlco/ Look-Both-Ways-and-No-Sink .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 146 citations
- Merging Statistical Feature via Adaptive Gate for Improved Text ClassificationXianming Li, Zongxi Li, Haoran Xie, Qing LiAAAI 2021 · 55 citations
- Modeling Instance Interactions for Joint Information Extraction with Neural High-Order Conditional Random FieldZixia Jia, Zhaohui Yan, Wenjuan Han, Zilong Zheng et al.ACL 2023 · 3 citations
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding ModelsChankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman et al.ICLR 2025
Related papers
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
- MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling CapabilitiesSavya Khosla, Aditi Tiwari, Kushal Kafle, Simon Jenni et al.ACL 2025
- When Attention Sink Emerges in Language Models: An Empirical ViewXiangming Gu, Tianyu Pang, Chao Du, Qian Liu et al.ICLR 2025
- Repetition Improves Language Model EmbeddingsJacob Mitchell Springer, Suhas Kotha, Daniel Fried, Graham Neubig et al.ICLR 2025
- Representation Deficiency in Masked Language ModelingYu Meng, Jitin Krishnan, Sinong Wang, Qifan Wang et al.ICLR 2024 · 12 citations
