Generalizable and Composable Multi-Model Embedding Translation
Beining Yang, Yang Cao
摘要
Embedding translation enables interoperability across embedding models, allowing embedding vectors to be reused without costly re-embedding. However, existing methods are typically evaluated under simplified pairwise and in-domain settings and behave as black boxes at inference time, leading to unreliable performance under out-of-distribution (OOD) inputs, multi-model mixing, and composed translations. We analyze embedding translation from a geometric perspective and derive an interpretable error bound that explains systematic error amplification under OOD inputs, mixing and chaining. Building on this, we propose a geometry-aware confidence metric and a Hierarchical Mixture of Experts (H-MoE) framework with localized, parameter-efficient adaptation. Following MTEB leaderboard, we conduct large-scale experiments over 10 embedding models and 6 benchmarks across 90 translation directions. H-MoE outperforms every baseline for every model pair over every benchmark under OOD scenarios. Furthermore, multi-model mixing and chaining only degrade our performance in Recall@100 by 0.5% ∼ 2.6%, compared to 7.2% ∼ 92.3% recall drop by existing methods. Code is available at https: //github.com/DBgroup-Edinburgh/ embedding-translation. Q1: Can we bound error 𝒆(𝒇 𝑨→𝑩 ) under Outof-Distribution (OOD)? Q2: Can we minimize gap 𝒆(𝒇 𝑨→𝑩 ; 𝒇 𝑪→𝑩 ) 𝒗𝒔. 𝒎𝒂𝒙(𝒆 𝒇 𝑨→𝑩 , 𝒆 𝒇 𝑪→𝑩 )? Q3: Can we minimize gap 𝒆(𝒇 𝐀→𝑪 ∘ 𝒇 𝐂→𝑩 ) 𝒗𝐬. 𝒆(𝒇 𝐀→𝑩
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li 等ICML 2020 · 被引用 193 次
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 被引用 69 次
- Fact or Fiction: Verifying Scientific ClaimsDavid Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang 等EMNLP 2020 · 被引用 6 次
- Integrating Vector Databases across Embedding ModelsBeining Yang, Yang Cao, Yang RenSIGMOD 2026 · 被引用 3 次
- Embedding-Converter: A Unified Framework for Cross-Model Embedding TransformationJinsung Yoon, Sercan Ö. ArikACL 2025 · 被引用 2 次
相关 Paper
- HUME: Measuring the Human-Model Performance Gap in Text Embedding TasksAdnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff 等ICLR 2026 · 被引用 8 次
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 · 被引用 27 次
- Geodesic Multi-Modal Mixup for Robust Fine-TuningChangdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim 等NeurIPS 2023 · 被引用 49 次
- MoE-RBench: Towards Building Reliable Language Models with Sparse Mixture-of-ExpertsGuanjie Chen, Xinyu Zhao, Tianlong Chen, Yu ChengICML 2024 · 被引用 8 次
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-ExpertsJingnan Gao, Zhe Wang, Xianze Fang, Xingyu Ren 等CVPR 2026 · 被引用 19 次
