FedVS: Towards Federated Vector Similarity Search with Filters
Zeheng Fan, Yuxiang Zeng, Zhuanglin Zheng, Binhan Yang, Yongxin Tong
Abstract
Vectors are used to represent unstructured data with their embeddings and associated attributes. Similarity search over large-scale vector datasets has gained significant interest from both industry and academia. It aims to identify the k nearest neighbors to a query object from vectors that satisfy a given attribute filter constraint. Despite its popularity, most solutions focus on single-sourced data and overlook the need for vector retrieval across federated datasets. To fill this gap, we introduce a new problem, federated vector similarity search with filters, which enables privacy-preserving vector retrieval over multi-sourced data held by mutually untrusted providers. While some solutions can be adapted, they struggle with low recall, excessive search latency, or high communication cost. To address these challenges, we propose FedVS, a privacy-preserving framework enhanced with indexing and pruning based on Trusted Execution Environment (TEE). We also provide a comprehensive theoretical analysis, including complexity, security, and approximation guarantees for recall. Moreover, we deploy our solution over real-world vector databases and conduct extensive experiments. The results demonstrate that our solution outperforms state-of-the-art methods in both effectiveness and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 539c1556-537c-471f-b774-484d2e533090Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
- SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood SearchQi Chen, Bing Zhao, Haidong Wang, Mingqin Li et al.NeurIPS 2021 · 219 citations
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 178 citations
- VBASE: Unifying Online Vector Similarity Search and Relational Queries via Relaxed MonotonicityQianxi Zhang, Shuotao Xu, Qi Chen, Guoxin Sui et al.OSDI 2023 · 75 citations
- Why Are Learned Indexes So Effective?Paolo Ferragina, Fabrizio Lillo, Giorgio VinciguerraICML 2020 · 63 citations
Related papers
- Federated Retrieval Over Embedding-Heterogeneous Vector DatabasesYuxiang Wang, Yongxin Tong, Zimu Zhou, Ziyuan He et al.ICDE 2026 · 1 citation
- Privacy-Preserving Approximate Nearest Neighbor Search on High-Dimensional DataYingfan Liu, Yandi Zhang, Jiadong Xie, Hui Li et al.ICDE 2025 · 2 citations
- CoTra: Towards Efficient and Scalable Distributed Vector Search with RDMAXiangyu Zhi, Meng Chen, Xiao Yan, Baotong Lu et al.SIGMOD 2026 · 7 citations
- RWalks: Random Walks as Attribute Diffusers for Filtered Vector SearchAnas Ait Aomar, Karima Echihabi, Marco Arnaboldi, Ioannis Alagiannis et al.SIGMOD 2025 · 4 citations
- VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis]Xiang Zhang, Chao Zhang, Ju Fan, Guoliang Li et al.SIGMOD 2026 · 5 citations
