Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix Trees
Arun Subramaniyan, Jack Wadden, Kush Goliya, Nathan Ozog, Xiao Wu, Satish Narayanasamy, David T. Blaauw, Reetuparna Das
摘要
Read alignment is a time-consuming step in genome sequencing analysis. The most widely used software for read alignment, BWA-MEM, and the recently published faster version BWA-MEM2 are based on the seed-and-extend paradigm for read alignment. The seeding step of read alignment is a major bottleneck contributing ∼40% to the overall execution time of BWA-MEM2 when aligning whole human genome reads from the Platinum Genomes dataset. This is because both BWA-MEM and BWA-MEM2 use a compressed index structure called the FMD-Index, which results in high bandwidth requirements, primarily due to its character-by-character processing of reads. For instance, to seed each read (101 DNA base-pairs stored in 37.8 bytes), the FMD-Index solution in BWA-MEM2 requires ∼68.5 KB of index data.
We propose a novel indexing data structure named Enumerated Radix Tree (ERT) and design a custom seeding accelerator based on it. ERT improves bandwidth efficiency of BWA-MEM2 by 4.5× while guaranteeing 100% identical output to the original software, and still fitting in 64 GB DRAM. Overall, the proposed seeding accelerator implemented on AWS F1 FPGA (f1.4xlarge) improves seeding throughput of BWA-MEM2 by 3.3×. When combined with seed-extension accelerators, we observe a 2.1× improvement in overall read alignment throughput over BWA-MEM2. The software implementation of ERT is integrated into BWA-MEM2 (ert branch: https://github.com/bwa-mem2/bwa-mem2/tree/ert) and is open sourced for the benefit of the research community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage ProcessingNika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer 等ISCA 2024 · 被引用 15 次
- TALCO: Tiling Genome Sequence Alignment Using Convergence of Traceback PointersSumit Walia, Cheng Ye, Arkid Bera, Dhruvi Lodhavia 等HPCA 2024 · 被引用 14 次
- CASA: An Energy-Efficient and High-Speed CAM-based SMEM Seeding Accelerator for Genome AlignmentYi Huang, Lingkun Kong, Dibei Chen, Zhiyu Chen 等MICRO 2023 · 被引用 5 次
- SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence AnalysisNika Mansouri-Ghiasi, Talu Güloglu, Harun Mustafa, Can Firtina 等HPCA 2026 · 被引用 3 次
它引用的顶会 Paper2
- SeedEx: A Genome Sequencing Accelerator for Optimal Alignments in Subminimal SpaceDaichi Fujiki, Shunhao Wu, Nathan Ozog, Kush Goliya 等MICRO 2020 · 被引用 52 次
- GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence AnalysisDamla Senol Cali, Gurpreet S. Kalsi, Zülal Bingöl, Can Firtina 等MICRO 2020 · 被引用 23 次
相关 Paper
- BLESS: Bandwidth and Locality Enhanced SMEM Seeding Acceleration for DNA SequencingSeunghee Han, Seungjae Moon, Teokkyu Suh, Jaehoon Heo 等ISCA 2024 · 被引用 5 次
- Faster and Cheaper: Pushing the Sequence Alignment Throughput with Commercial CPUsZhonghai Zhang, Yewen Li, Ke Meng, Chunming Zhang 等PPoPP 2026 · 被引用 1 次
- Lembas: Cost-Efficient Genome Alignment with External Memory and FPGA AccelerationSeongyoung Kang, Se-Min Lim, Sang-Woo JunISCA 2026
- NvWa: Enhancing Sequence Alignment Accelerator Throughput via Hardware SchedulingYewen Li, Xueqi Li, Ruihao Gao, Wanqi Liu 等HPCA 2023 · 被引用 6 次
- GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read MappingJulien Eudine, Chu Li, Zhuo Cheng, Renzo Andri 等HPCA 2026 · 被引用 2 次
