IDentity with Locality: An Ideal Hash for Gene Sequence Search
Tianyi Zhang, Gaurav Gupta, Aditya Desai, Anshumali Shrivastava
Abstract
Gene sequence search is a fundamental operation in computational genomics with broad applications in medicine, evolutionary biology, metagenomics, and more. Due to the petabyte scale of genome archives, most gene search systems now use hashing-based data structures such as Bloom Filters (BF). The state-of-the-art systems such as Compact bit-slicing signature index (COBS) [4] and Repeated And Merged Bloom filters (RAMBO) [21] use BF with Random Hash (RH) functions for gene representation and identification. The standard recipe is to cast the gene search problem as a sequence of membership problems testing if each subsequent gene substring (called kmer) of 𝑄 is present in the set of kmers of the entire gene database 𝐷. We observe that RH functions, which are crucial to the memory and the computational advantage of BF, are also detrimental to the system performance of gene-search systems. While subsequent kmers being queried are likely very similar, RH, oblivious to any similarity, uniformly distributes the kmers to different parts of potentially large BF, thus triggering excessive cache misses and causing system slowdown. We propose a novel hash function called the Identity with Locality (IDL) hash family, which co-locates the keys close in input space without causing collisions. This approach ensures both cache locality and key preservation. IDL functions can be a drop-in replacement for RH functions and help improve the performance of information retrieval systems. We give a simple but practical construction of IDL function families and show that replacing the RH with IDL functions reduces cache misses by a factor of 5×, thus improving query and indexing times of SOTA methods such as COBS and RAMBO by factors up to 2× without compromising their quality. We also provide a theoretical analysis of the false positive rate of BF with IDL functions. Our hash function is the first study that bridges Locality Sensitive Hash (LSH) and RH to obtain cache efficiency. Our design and analysis could be of independent theoretical interest.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b68d5a6-e9d8-4a08-a08e-a7fce2f35976Builds on2
- Fast Processing and Querying of 170TB of Genomics Data via a Repeated And Merged BloOm Filter (RAMBO)Gaurav Gupta, Minghao Yan, Benjamin Coleman, Bryce Kille et al.SIGMOD 2021 · 19 citations
- BLISS: A Billion scale Index using Iterative Re-partitioningGaurav Gupta, Tharun Medini, Anshumali Shrivastava, Alexander J. SmolaKDD 2022 · 14 citations
Related papers
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim et al.ISCA 2022 · 66 citations
- The next 50 Years in Database Indexing or: The Case for Automatically Generated Index StructuresJens Dittrich, Joris Nix, Christian SchönVLDB 2022 · 12 citations
- Locality Sensitive Hashing for Optimizing Subgraph Query Processing in Parallel Computing SystemsPeng Peng, Shengyi Ji, Zhen Tian, Hongbo Jiang et al.KDD 2023 · 1 citation
- BLESS: Bandwidth and Locality Enhanced SMEM Seeding Acceleration for DNA SequencingSeunghee Han, Seungjae Moon, Teokkyu Suh, Jaehoon Heo et al.ISCA 2024 · 5 citations
- EDIndex: Enabling Fast Data Queries in Edge Storage SystemsQiang He, Siyu Tan, Feifei Chen, Xiaolong Xu et al.SIGIR 2023 · 34 citations
