Integrating Vector Databases across Embedding Models
Beining Yang, Yang Cao, Yang Ren
Abstract
Vector databases have been widely used to implement similarity search over unstructured objects, e.g., documents and images. Each vector database is produced by an embedding model that encodes the objects in a way such that more similar objects are embedded to closer vectors, allowing us to use top-k vector search as an implementation of top-k object similarity search. It is common practice that different vector databases use distinct embedding models and the same object may be encoded by different embedding vectors across databases. As a result, one cannot share and integrate vector databases to expand similarity search across datasets, a property we take for granted for relational databases. In this work, we attempt to break the barrier between different vector databases, by developing an approach to integrating vector databases generated by different embedding models, with neither any access to the encoded data objects nor knowledge of the embedding models. Our approach is rooted in the local isometry hypothesis, a finding made via extensive experiments on real-life embedding vectors, and is backed up by theoretical analysis that bounds the quality of integrated vector database. Experimental results show that we can integrate vector databases produced by various popular embedding models, e.g., NV-embed-V2, OpenAI Ada, GloVe, Mistral and FastText, while offering high recall of top-k similarity search over the integrated datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a63b9ab4-5d2b-4111-98bd-81e46760bae1Cited by top-tier papers3
- Generalizable and Composable Multi-Model Embedding TranslationBeining Yang, Yang CaoICML 2026
- Vector Linking via Cross-Model Local Isometric ConsistencyZiying Chen, Yang Cao, He Sun, Beining Yang et al.ICML 2026
- FedAugment: Table Augmentation Search over Decentralized Data RepositoriesLennart Behme, Emil Badura, Leonard Geißler, Matthias Böhm et al.VLDB 2026
Related papers
- SQLVec: SQL-Based Vector Similarity SearchZhequn Zhang, Yuanyuan Zhu, Hao Zhang, Jeffrey Xu YuICDE 2026
- Federated Retrieval Over Embedding-Heterogeneous Vector DatabasesYuxiang Wang, Yongxin Tong, Zimu Zhou, Ziyuan He et al.ICDE 2026 · 1 citation
- FedVS: Towards Federated Vector Similarity Search with FiltersZeheng Fan, Yuxiang Zeng, Zhuanglin Zheng, Binhan Yang et al.KDD 2025 · 1 citation
- On the Theoretical Limitations of Embedding-Based RetrievalOrion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk LeeICLR 2026 · 138 citations
- BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector DatabasesGuoxin Kang, Zhongxin Ge, Jingpei Hu, Xueya Zhang et al.VLDB 2025 · 4 citations
