Repurposing Language Models into Embedding Models: Finding the Compute-Optimal Recipe
Albert Q. Jiang, Alicja Ziarko, Bartosz Piotrowski, Wenda Li, Mateja Jamnik, Piotr Milos
Abstract
Text embeddings are essential for many tasks, such as document retrieval, clustering, and semantic similarity assessment. In this paper, we study how to contrastively train text embedding models in a compute-optimal fashion, given a suite of pre-trained decoder-only language models. Our innovation is an algorithm that produces optimal configurations of model sizes, data quantities, and fine-tuning methods for text-embedding models at different computational budget levels. The resulting recipe, which we obtain through extensive experiments, can be used by practitioners to make informed design choices for their embedding models. Specifically, our findings suggest that full fine-tuning and low-rank adaptation fine-tuning produce optimal models at lower and higher computational budgets respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e9def12-8a9f-4a66-8f2e-48dc67a54d18Cited by top-tier papers4
- Contrastive Representations for Temporal ReasoningAlicja Ziarko, Michal Bortkiewicz, Michal Zawalski, Benjamin Eysenbach et al.NeurIPS 2025 · 8 citations
- Attention with Trained Embeddings Provably Selects Important TokensDiyuan Wu, Aleksandr Shevchenko, Samet Oymak, Marco MondelliNeurIPS 2025
- ClarifyVC: Clarifying Ambiguous Commands in Vehicle Control with a Hybrid Data Augmentation PipelineHange Zhou, Zhonglin Jiang, yingjie cui, Mingzhe Zhang et al.ICLR 2026
- Enhancing Lexicon-Based Text Embeddings with Large Language ModelsYibin Lei, Tao Shen, Yu Cao, Andrew YatesACL 2025
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Improving Text Embeddings with Large Language ModelsLiang Wang, Nan Yang, Xiaolong Huang, Linjun Yang et al.ACL 2024
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 25 citations
- Effective post-training embedding compression via temperature control in contrastive trainingGeorgiana Dinu, Corey D. Barrett, Yi Xiang, Miguel Romero Calvo et al.ICLR 2025
- Training compute-optimal transformer encoder modelsMegi Dervishi, Alexandre Allauzen, Gabriel Synnaeve, Yann LeCunEMNLP 2025 · 1 citation
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
