Random-Access Ranked Retrieval and Similarity Search
Mohsen Dehghankar, Abolfazl Asudeh, Raghav Mittal, Suraj Shetiya, Gautam Das
Abstract
We extend Random Access, a fundamental operation that enables efficient search and exploration algorithms, to the modern interactive data systems based on Ranked Retrieval and Similarity Search, where orderings are dynamically defined over a high-dimensional feature space. This extension enables efficient solutions for a wide range of applications, from data analytics tools and database systems to recommendation systems and machine learning.
We formalize the Random-Access Ranked Retrieval (RAR) problem, and extend it to Similarity Search. Our algorithmic innovations include the development of a theoretically efficient algorithm based on geometric arrangements, achieving logarithmic query time. However, this method suffers from exponential space complexity in high dimensions. Therefore, we develop a second class of algorithms based on 𝜀-sampling, which consume a linear space. Since exactly locating the tuple at a specific rank is challenging due to its connection to the range counting problem, we introduce a relaxed variant called 𝜅-Random-Access Ranked Retrieval, which returns a small subset of size 𝜅 guaranteed to contain the target tuple. To solve this problem efficiently, we define an intermediate problem, Stripe Range Retrieval (SRR), and design a hierarchical sampling data structure tailored for narrow stripe range queries. Our method achieves practical scalability in both data size and dimensionality. We prove near-optimal bounds on the efficiency of our algorithms and validate their performance through extensive experiments on real and synthetic datasets, demonstrating scalability to millions of tuples and hundreds of dimensions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor SearchMengzhao Wang, Xiaoliang Xu, Qiang Yue, Yuxiang WangVLDB 2021 · 354 citations
- JENNER: Just-in-time Enrichment in Query ProcessingDhrubajyoti Ghosh, Peeyush Gupta, Sharad Mehrotra, Roberto Yus et al.VLDB 2022 · 5 citations
- Approximation-First Timeseries Monitoring Query At ScaleZeying Zhu, Jonathan Chamberlain, Kenny Wu, David Starobinski et al.VLDB 2025 · 2 citations
Related papers
- Retrieval with Learned SimilaritiesBailu Ding, Jiaqi ZhaiWWW 2025 · 3 citations
- ProMIPS: Efficient High-Dimensional c-Approximate Maximum Inner Product Search with a Lightweight IndexYang Song, Yu Gu, Rui Zhang, Ge YuICDE 2021 · 16 citations
- Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random ForestsChristian Lülf, Denis Mayr Lima Martins, Marcos Antonio Vaz Salles, Yongluan Zhou et al.VLDB 2023 · 5 citations
- MinSearch: An Efficient Algorithm for Similarity Search under Edit DistanceHaoyu Zhang, Qin ZhangKDD 2020 · 11 citations
- Lazy Search TreesBryce Sandlund, Sebastian WildFOCS 2020 · 1 citation
