Look Both Ways and No Sink: Converting LLMs into Text Encoders without Training
Ziyong Lin, Haoyi Wu, Shu Wang, Kewei Tu, Zilong Zheng, Zixia Jia
摘要
Recent advancements have demonstrated the advantage of converting pretrained large language models into powerful text encoders by enabling bidirectional attention in transformer layers. However, existing methods often require extensive training on large-scale datasets, posing challenges in low-resource, domain-specific scenarios. In this work, we show that a pretrained large language model can be converted into a strong text encoder without additional training. We first conduct a comprehensive empirical study to investigate different conversion strategies and identify the impact of the attention sink phenomenon on the performance of converted encoder models. Based on our findings, we propose a novel approach that enables bidirectional attention and suppresses the attention sink phenomenon, resulting in superior performance. Extensive experiments on multiple domains demonstrate the effectiveness of our approach. Our work provides new insights into the training-free conversion of text encoders in low-resource scenarios and contributes to the advancement of domain-specific text representation generation. Our code is available at https://github.com/bigai-nlco/ Look-Both-Ways-and-No-Sink .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 被引用 146 次
- Merging Statistical Feature via Adaptive Gate for Improved Text ClassificationXianming Li, Zongxi Li, Haoran Xie, Qing LiAAAI 2021 · 被引用 55 次
- Modeling Instance Interactions for Joint Information Extraction with Neural High-Order Conditional Random FieldZixia Jia, Zhaohui Yan, Wenjuan Han, Zilong Zheng 等ACL 2023 · 被引用 3 次
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding ModelsChankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman 等ICLR 2025
相关 Paper
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
- MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling CapabilitiesSavya Khosla, Aditi Tiwari, Kushal Kafle, Simon Jenni 等ACL 2025
- When Attention Sink Emerges in Language Models: An Empirical ViewXiangming Gu, Tianyu Pang, Chao Du, Qian Liu 等ICLR 2025
- Repetition Improves Language Model EmbeddingsJacob Mitchell Springer, Suhas Kotha, Daniel Fried, Graham Neubig 等ICLR 2025
- Representation Deficiency in Masked Language ModelingYu Meng, Jitin Krishnan, Sinong Wang, Qifan Wang 等ICLR 2024 · 被引用 12 次
