Endeavor: Efficient PairHMM for Detection of DNA Variants in Genome-Scale Datasets
Miguel Graça, Aleksandar Ilic
摘要
DNA variant calling represents a key operation in bioinformatics pipelines that aims at identifying genetic variants. Given an evidenced explosion in genomic data availability, there is an urgent need for a high-performant, portable and efficient solution for variant calling, which can further improve our understanding of genomic structure and genetic basis for complex diseases. In its most common formulation, the Pair Hidden Markov Model (PairHMM) algorithm for variant calling stands as the main bottleneck in the pipeline, accounting for up to 70% of the execution time in large-scale genomic datasets. The state-of-the-art approaches for accelerating PairHMM in CPUs and GPUs do not scale to long DNA sequences and only explore very limited anti-diagonal data parallelism, which yields poor performance. In this work, Endeavor is proposed as a new parallelization strategy for PairHMM that redefines its traditional formulation to explore row-level fine-grained parallelism without loss in solution accuracy. Based on this, a novel and portable SIMD-based approach is derived for efficient and high-performance processing of short and long sequences in CPUs and GPUs, leveraging novel levels of parallelism and synchronization to achieve high throughput in sequences up to 100k basepairs for the first time. Evaluation on Intel and AMD CPUs shows that Endeavor outperforms GKL up to 2.14x in peak throughput and GATK HaplotypeCaller by at least 2x in real-world datasets, while NVIDIA and AMD GPUs achieve up to 2.05x speedups in genome-scale datasets when compared to state-of-the-art GPU-based methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read MappingSeongyeon Park, Junguk Hong, Jaeyong Song, Hajin Kim 等PPoPP 2024 · 被引用 7 次
- NvWa: Enhancing Sequence Alignment Accelerator Throughput via Hardware SchedulingYewen Li, Xueqi Li, Ruihao Gao, Wanqi Liu 等HPCA 2023 · 被引用 6 次
- GenPIP: In-Memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read MappingHaiyu Mao, Mohammed Alser, Mohammad Sadrosadati, Can Firtina 等MICRO 2022 · 被引用 32 次
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim 等ISCA 2022 · 被引用 66 次
- RapidGKC: GPU-Accelerated K-Mer CountingYiran Cheng, Xibo Sun, Qiong LuoICDE 2024 · 被引用 4 次
