Robust Tree-based Learned Vector Index with Query-aware Repartitioning
Wenqing Wei, Defu Lian, Qingshuai Feng, Yongji Wu
Abstract
Approximate Vector Retrieval (AVR), which aims to efficiently retrieve the most similar items from a large dataset, is a fundamental task in a variety of applications such as information retrieval, recommender systems, and large language models. Advances in representation learning and multimodal neural models have enabled diverse data types (e.g., text, images, audio) to be embedded into a shared vector space, facilitating similarity-based retrieval in AVR. While single-modal AVR assumes query and database embeddings follow the same distribution (In-Distribution, ID), cross-modal AVR introduces a distribution shift, where query vectors (e.g., text) are Out-of-Distribution (OOD) relative to the database (e.g., images). This mismatch complicates retrieval and degrades accuracy, making it a key challenge in AVR. Existing methods typically focus on either ID or OOD queries but struggle to handle both within a unified framework.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- RoarGraph: A Projected Bipartite Graph for Efficient Cross-Modal Approximate Nearest Neighbor SearchMeng Chen, Kai Zhang, Zhenying He, Yinan Jing et al.VLDB 2024 · 27 citations
- Incomplete Cross-modal Retrieval with Dual-Aligned Variational AutoencodersMengmeng Jing, Jingjing Li, Lei Zhu, Ke Lu et al.ACM MM 2020 · 63 citations
- Universal Vision-Language Dense Retrieval: Learning A Unified Representation Space for Multi-Modal RetrievalZhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu et al.ICLR 2023 · 6 citations
- Efficient and High-Fidelity Omni Modality RetrievalChuong Huynh, Manh Luong, Abhinav ShrivastavaCVPR 2026 · 3 citations
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
