IDentity with Locality: An Ideal Hash for Gene Sequence Search
Tianyi Zhang, Gaurav Gupta, Aditya Desai, Anshumali Shrivastava
摘要
Gene sequence search is a fundamental operation in computational genomics with broad applications in medicine, evolutionary biology, metagenomics, and more. Due to the petabyte scale of genome archives, most gene search systems now use hashing-based data structures such as Bloom Filters (BF). The state-of-the-art systems such as Compact bit-slicing signature index (COBS) [4] and Repeated And Merged Bloom filters (RAMBO) [21] use BF with Random Hash (RH) functions for gene representation and identification. The standard recipe is to cast the gene search problem as a sequence of membership problems testing if each subsequent gene substring (called kmer) of 𝑄 is present in the set of kmers of the entire gene database 𝐷. We observe that RH functions, which are crucial to the memory and the computational advantage of BF, are also detrimental to the system performance of gene-search systems. While subsequent kmers being queried are likely very similar, RH, oblivious to any similarity, uniformly distributes the kmers to different parts of potentially large BF, thus triggering excessive cache misses and causing system slowdown. We propose a novel hash function called the Identity with Locality (IDL) hash family, which co-locates the keys close in input space without causing collisions. This approach ensures both cache locality and key preservation. IDL functions can be a drop-in replacement for RH functions and help improve the performance of information retrieval systems. We give a simple but practical construction of IDL function families and show that replacing the RH with IDL functions reduces cache misses by a factor of 5×, thus improving query and indexing times of SOTA methods such as COBS and RAMBO by factors up to 2× without compromising their quality. We also provide a theoretical analysis of the false positive rate of BF with IDL functions. Our hash function is the first study that bridges Locality Sensitive Hash (LSH) and RH to obtain cache efficiency. Our design and analysis could be of independent theoretical interest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Fast Processing and Querying of 170TB of Genomics Data via a Repeated And Merged BloOm Filter (RAMBO)Gaurav Gupta, Minghao Yan, Benjamin Coleman, Bryce Kille 等SIGMOD 2021 · 被引用 19 次
- BLISS: A Billion scale Index using Iterative Re-partitioningGaurav Gupta, Tharun Medini, Anshumali Shrivastava, Alexander J. SmolaKDD 2022 · 被引用 14 次
相关 Paper
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim 等ISCA 2022 · 被引用 66 次
- The next 50 Years in Database Indexing or: The Case for Automatically Generated Index StructuresJens Dittrich, Joris Nix, Christian SchönVLDB 2022 · 被引用 12 次
- Locality Sensitive Hashing for Optimizing Subgraph Query Processing in Parallel Computing SystemsPeng Peng, Shengyi Ji, Zhen Tian, Hongbo Jiang 等KDD 2023 · 被引用 1 次
- BLESS: Bandwidth and Locality Enhanced SMEM Seeding Acceleration for DNA SequencingSeunghee Han, Seungjae Moon, Teokkyu Suh, Jaehoon Heo 等ISCA 2024 · 被引用 5 次
- EDIndex: Enabling Fast Data Queries in Edge Storage SystemsQiang He, Siyu Tan, Feifei Chen, Xiaolong Xu 等SIGIR 2023 · 被引用 34 次
