Embedding-Converter: A Unified Framework for Cross-Model Embedding Transformation
Jinsung Yoon, Sercan Ö. Arik
摘要
Embedding models play a crucial role in machine learning. However, the continuous development of new models presents a major challenge: migrating to a potentially superior model often requires the computationally expensive process of re-embedding entire datasets-without any guarantee of performance improvement. This paper presents Embedding-Converter, a novel framework for efficiently transforming embeddings between different models, thus avoiding costly 'reembedding'. The proposed approach achieves 100 times faster and cheaper computations in real-world applications. Experiments show that Embedding-Converter not only streamlines transitions to new models, but can also improve upon the source model's performance, approaching that of the target model. This facilitates efficient evaluation and broader adoption of new embedding models by significantly reducing the overhead of model switching. Furthermore, Embedding-Converter addresses latency limitations by enabling the use of smaller models for online tasks while still benefiting from the performance of larger models offline. By promoting the release of converters alongside new embedding models, Embedding-Converter fosters a more dynamic and accessible ecosystem for embedding model development and deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 被引用 69 次
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned RepresentationsRobin Vujanic, Thomas RückstießACL 2026 · 被引用 5 次
- Generalizable and Composable Multi-Model Embedding TranslationBeining Yang, Yang CaoICML 2026
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
- Forward Compatible Training for Large-Scale Embedding Retrieval SystemsVivek Ramanujan, Pavan Kumar Anasosalu Vasu, Ali Farhadi, Oncel Tuzel 等CVPR 2022 · 被引用 12 次
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding ModelsChankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman 等ICLR 2025
相关 Paper
- Drift-Adapter: A Practical Approach to Near Zero-Downtime Embedding Model Upgrades in Vector DatabasesHarshil VejendlaEMNLP 2025 · 被引用 1 次
- Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding DimensionsJinsung Yoon, Rajarishi Sinha, Sercan Ömer Arik, Tomas PfisterEMNLP 2024 · 被引用 1 次
- EmbedLLM: Learning Compact Representations of Large Language ModelsRichard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li 等ICLR 2025
- InferDB: In-Database Machine Learning Inference Using IndexesRicardo Salazar-Díaz, Boris Glavic, Tilmann RablVLDB 2024 · 被引用 13 次
- Experimental Analysis of Large-scale Learnable Vector Storage CompressionHailin Zhang, Penghao Zhao, Xupeng Miao, Yingxia Shao 等VLDB 2024 · 被引用 20 次
