Endeavor: Efficient PairHMM for Detection of DNA Variants in Genome-Scale Datasets
Miguel Graça, Aleksandar Ilic
Abstract
DNA variant calling represents a key operation in bioinformatics pipelines that aims at identifying genetic variants. Given an evidenced explosion in genomic data availability, there is an urgent need for a high-performant, portable and efficient solution for variant calling, which can further improve our understanding of genomic structure and genetic basis for complex diseases. In its most common formulation, the Pair Hidden Markov Model (PairHMM) algorithm for variant calling stands as the main bottleneck in the pipeline, accounting for up to 70% of the execution time in large-scale genomic datasets. The state-of-the-art approaches for accelerating PairHMM in CPUs and GPUs do not scale to long DNA sequences and only explore very limited anti-diagonal data parallelism, which yields poor performance. In this work, Endeavor is proposed as a new parallelization strategy for PairHMM that redefines its traditional formulation to explore row-level fine-grained parallelism without loss in solution accuracy. Based on this, a novel and portable SIMD-based approach is derived for efficient and high-performance processing of short and long sequences in CPUs and GPUs, leveraging novel levels of parallelism and synchronization to achieve high throughput in sequences up to 100k basepairs for the first time. Evaluation on Intel and AMD CPUs shows that Endeavor outperforms GKL up to 2.14x in peak throughput and GATK HaplotypeCaller by at least 2x in real-world datasets, while NVIDIA and AMD GPUs achieve up to 2.05x speedups in genome-scale datasets when compared to state-of-the-art GPU-based methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f79ebbea-2b32-4ee8-ab2b-0672a03c7a2eBuilds on1
Related papers
- AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read MappingSeongyeon Park, Junguk Hong, Jaeyong Song, Hajin Kim et al.PPoPP 2024 · 7 citations
- NvWa: Enhancing Sequence Alignment Accelerator Throughput via Hardware SchedulingYewen Li, Xueqi Li, Ruihao Gao, Wanqi Liu et al.HPCA 2023 · 6 citations
- GenPIP: In-Memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read MappingHaiyu Mao, Mohammed Alser, Mohammad Sadrosadati, Can Firtina et al.MICRO 2022 · 32 citations
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim et al.ISCA 2022 · 66 citations
- RapidGKC: GPU-Accelerated K-Mer CountingYiran Cheng, Xibo Sun, Qiong LuoICDE 2024 · 4 citations
