Accelerated Seeding for Genome Sequence Alignment with Enumerated Radix Trees
Arun Subramaniyan, Jack Wadden, Kush Goliya, Nathan Ozog, Xiao Wu, Satish Narayanasamy, David T. Blaauw, Reetuparna Das
Abstract
Read alignment is a time-consuming step in genome sequencing analysis. The most widely used software for read alignment, BWA-MEM, and the recently published faster version BWA-MEM2 are based on the seed-and-extend paradigm for read alignment. The seeding step of read alignment is a major bottleneck contributing ∼40% to the overall execution time of BWA-MEM2 when aligning whole human genome reads from the Platinum Genomes dataset. This is because both BWA-MEM and BWA-MEM2 use a compressed index structure called the FMD-Index, which results in high bandwidth requirements, primarily due to its character-by-character processing of reads. For instance, to seed each read (101 DNA base-pairs stored in 37.8 bytes), the FMD-Index solution in BWA-MEM2 requires ∼68.5 KB of index data.
We propose a novel indexing data structure named Enumerated Radix Tree (ERT) and design a custom seeding accelerator based on it. ERT improves bandwidth efficiency of BWA-MEM2 by 4.5× while guaranteeing 100% identical output to the original software, and still fitting in 64 GB DRAM. Overall, the proposed seeding accelerator implemented on AWS F1 FPGA (f1.4xlarge) improves seeding throughput of BWA-MEM2 by 3.3×. When combined with seed-extension accelerators, we observe a 2.1× improvement in overall read alignment throughput over BWA-MEM2. The software implementation of ERT is integrated into BWA-MEM2 (ert branch: https://github.com/bwa-mem2/bwa-mem2/tree/ert) and is open sourced for the benefit of the research community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c67b8fc4-9c0f-4642-9bea-3aa094ecd0c0Cited by top-tier papers4
- MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage ProcessingNika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer et al.ISCA 2024 · 15 citations
- TALCO: Tiling Genome Sequence Alignment Using Convergence of Traceback PointersSumit Walia, Cheng Ye, Arkid Bera, Dhruvi Lodhavia et al.HPCA 2024 · 14 citations
- CASA: An Energy-Efficient and High-Speed CAM-based SMEM Seeding Accelerator for Genome AlignmentYi Huang, Lingkun Kong, Dibei Chen, Zhiyu Chen et al.MICRO 2023 · 5 citations
- SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence AnalysisNika Mansouri-Ghiasi, Talu Güloglu, Harun Mustafa, Can Firtina et al.HPCA 2026 · 3 citations
Builds on2
- SeedEx: A Genome Sequencing Accelerator for Optimal Alignments in Subminimal SpaceDaichi Fujiki, Shunhao Wu, Nathan Ozog, Kush Goliya et al.MICRO 2020 · 52 citations
- GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence AnalysisDamla Senol Cali, Gurpreet S. Kalsi, Zülal Bingöl, Can Firtina et al.MICRO 2020 · 23 citations
Related papers
- BLESS: Bandwidth and Locality Enhanced SMEM Seeding Acceleration for DNA SequencingSeunghee Han, Seungjae Moon, Teokkyu Suh, Jaehoon Heo et al.ISCA 2024 · 5 citations
- Faster and Cheaper: Pushing the Sequence Alignment Throughput with Commercial CPUsZhonghai Zhang, Yewen Li, Ke Meng, Chunming Zhang et al.PPoPP 2026 · 1 citation
- Lembas: Cost-Efficient Genome Alignment with External Memory and FPGA AccelerationSeongyoung Kang, Se-Min Lim, Sang-Woo JunISCA 2026
- NvWa: Enhancing Sequence Alignment Accelerator Throughput via Hardware SchedulingYewen Li, Xueqi Li, Ruihao Gao, Wanqi Liu et al.HPCA 2023 · 6 citations
- GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read MappingJulien Eudine, Chu Li, Zhuo Cheng, Renzo Andri et al.HPCA 2026 · 2 citations
