Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions
Jinsung Yoon, Rajarishi Sinha, Sercan Ömer Arik, Tomas Pfister
Abstract
Embeddings from Large Language Models (LLMs) have emerged as critical components in various applications, particularly for information retrieval. While high-dimensional embeddings generally demonstrate superior performance as they contain more salient information, their practical application is frequently hindered by elevated computational latency and the associated higher cost. To address these challenges, we propose Matryoshka-Adaptor, a novel tuning framework designed for the customization of LLM embeddings. Matryoshka-Adaptor facilitates substantial dimensionality reduction while maintaining comparable performance levels, thereby achieving a significant enhancement in computational efficiency and costeffectiveness. Our framework directly modifies the embeddings from pre-trained LLMs which is designed to be seamlessly integrated with any LLM architecture, encompassing those accessible exclusively through blackbox APIs. Also, it exhibits efficacy in both unsupervised and supervised learning settings. A rigorous evaluation conducted across a diverse corpus of English, multilingual, and multimodal datasets consistently reveals substantial gains with Matryoshka-Adaptor. Notably, with Google and OpenAI Embedding APIs, Matryoshka-Adaptor achieves a reduction in dimensionality ranging from twoto twelve-fold without compromising performance across multiple BEIR datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f596b1f6-5153-4c74-9d8b-d87753c9c8e7Cited by top-tier papers4
- CSRv2: Unlocking Ultra-Sparse EmbeddingsLixuan Guo, Yifei Wang, Tiansheng Wen, Yifan Wang et al.ICLR 2026 · 7 citations
- ConceptCarve: Dynamic Realization of EvidenceEylon Caplan, Dan GoldwasserACL 2025
- SMEC:Rethinking Matryoshka Representation Learning for Retrieval Embedding CompressionBiao Zhang, Lixin Chen, Tong Liu, Bo ZhengEMNLP 2025
- ML-Embed: Inclusive and Efficient Embeddings for a Multilingual WorldZiyin Zhang, Zihan Liao, Hang Yu, Peng Di et al.ICML 2026
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Promptagator: Few-shot Dense Retrieval From 8 ExamplesZhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan et al.ICLR 2023 · 46 citations
- Generative Representational Instruction TuningNiklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang et al.ICLR 2025
Related papers
- Search-Adaptor: Embedding Customization for Information RetrievalJinsung Yoon, Yanfei Chen, Sercan Ö. Arik, Tomas PfisterACL 2024
- Fitting Into Any Shape: A Flexible LLM-Based Re-Ranker With Configurable Depth and WidthZheng Liu, Chaofan Li, Shitao Xiao, Chaozhuo Li et al.WWW 2025 · 1 citation
- EmbedLLM: Learning Compact Representations of Large Language ModelsRichard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li et al.ICLR 2025
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text EmbeddingsSonghao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui et al.KDD 2026 · 1 citation
- Enhancing Lexicon-Based Text Embeddings with Large Language ModelsYibin Lei, Tao Shen, Yu Cao, Andrew YatesACL 2025
