Generalizable and Composable Multi-Model Embedding Translation
Beining Yang, Yang Cao
Abstract
Embedding translation enables interoperability across embedding models, allowing embedding vectors to be reused without costly re-embedding. However, existing methods are typically evaluated under simplified pairwise and in-domain settings and behave as black boxes at inference time, leading to unreliable performance under out-of-distribution (OOD) inputs, multi-model mixing, and composed translations. We analyze embedding translation from a geometric perspective and derive an interpretable error bound that explains systematic error amplification under OOD inputs, mixing and chaining. Building on this, we propose a geometry-aware confidence metric and a Hierarchical Mixture of Experts (H-MoE) framework with localized, parameter-efficient adaptation. Following MTEB leaderboard, we conduct large-scale experiments over 10 embedding models and 6 benchmarks across 90 translation directions. H-MoE outperforms every baseline for every model pair over every benchmark under OOD scenarios. Furthermore, multi-model mixing and chaining only degrade our performance in Recall@100 by 0.5% โผ 2.6%, compared to 7.2% โผ 92.3% recall drop by existing methods. Code is available at https: //github.com/DBgroup-Edinburgh/ embedding-translation. Q1: Can we bound error ๐(๐ ๐จโ๐ฉ ) under Outof-Distribution (OOD)? Q2: Can we minimize gap ๐(๐ ๐จโ๐ฉ ; ๐ ๐ชโ๐ฉ ) ๐๐. ๐๐๐(๐ ๐ ๐จโ๐ฉ , ๐ ๐ ๐ชโ๐ฉ )? Q3: Can we minimize gap ๐(๐ ๐โ๐ช โ ๐ ๐โ๐ฉ ) ๐๐ฌ. ๐(๐ ๐โ๐ฉ
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 088edc18-6365-4f94-9195-c6e7cd6bcf7cBuilds on6
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li et al.ICML 2020 ยท 193 citations
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 ยท 69 citations
- Fact or Fiction: Verifying Scientific ClaimsDavid Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang et al.EMNLP 2020 ยท 6 citations
- Integrating Vector Databases across Embedding ModelsBeining Yang, Yang Cao, Yang RenSIGMOD 2026 ยท 3 citations
- Embedding-Converter: A Unified Framework for Cross-Model Embedding TransformationJinsung Yoon, Sercan ร. ArikACL 2025 ยท 2 citations
Related papers
- HUME: Measuring the Human-Model Performance Gap in Text Embedding TasksAdnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff et al.ICLR 2026 ยท 8 citations
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 ยท 27 citations
- Geodesic Multi-Modal Mixup for Robust Fine-TuningChangdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim et al.NeurIPS 2023 ยท 49 citations
- MoE-RBench: Towards Building Reliable Language Models with Sparse Mixture-of-ExpertsGuanjie Chen, Xinyu Zhao, Tianlong Chen, Yu ChengICML 2024 ยท 8 citations
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-ExpertsJingnan Gao, Zhe Wang, Xianze Fang, Xingyu Ren et al.CVPR 2026 ยท 19 citations
