Lune

HPCA2026Top-tier venue

GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read Mapping

Julien Eudine, Chu Li, Zhuo Cheng, Renzo Andri, Can Firtina, Mohammad Sadrosadati, Nika Mansouri-Ghiasi, Konstantina Koliogeorgi, Anirban Nag, Arash Tavakkol, Haiyu Mao, Onur Mutlu

2026Year
2Citations

Abstract

Genome sequencing has become a central focus in computational biology due to its critical role in applications such as personalized medicine, disease outbreak tracking, and evolutionary research. A genome study typically begins with sequencing, which produces millions to billions of short DNA fragments known as reads. Extracting meaningful biological insights from these reads requires a computationally intensive step called read mapping, where each read is aligned to a reference genome. Read mapping for short reads comes in two forms: single-end and paired-end, with the latter being more prevalent due to its higher accuracy and support for advanced analysis. Read mapping remains a major performance bottleneck in genome analysis as a result of the extensive use of computationally intensive dynamic programming. Prior efforts have attempted to mitigate this cost by employing filters to identify and potentially discard computationally expensive matches and leveraging hardware accelerators to speed up the computations. While partially effective, these approaches have limitations. In particular, existing filters are often ineffective for paired-end reads, as they evaluate each read independently and exhibit relatively low filtering ratios.

In this work, we propose GenPairX, a hardware-algorithm codesigned accelerator that efficiently minimizes the computational load of paired-end read mapping while enhancing the throughput of memory-intensive operations. GenPairX introduces: (1) a novel filtering algorithm that jointly considers both reads in a pair to improve filtering effectiveness, and a lightweight alignment algorithm to replace most of the computationally expensive dynamic programming operations, and (2) two specialized hardware mechanisms to support the proposed algorithms. The proposed hardware addresses the high memory bandwidth demands of the read filtering process via orchestration of memory accesses over high-bandwidth memory channels, and accelerates the alignment of candidate reads via simple vectorized logical XOR operators. Our evaluations show that GenPairX delivers substantial performance improvements over state-of-the-art solutions, achieving 1575× and 1.43× higher throughput per watt compared to leading CPU-based and accelerator-based read mappers, respectively, all without compromising accuracy.

We carefully design a high-throughput, balanced end-to-end read mapping pipeline and evaluate the impact of the proposed algorithmic and hardware contributions on the end-to-end performance of paired-end read mapping. We integrate GenPairX with GenDP [78], a hardware accelerator that serves as a fallback mechanism for the small fraction of read-pairs that cannot be mapped or aligned by GenPairX. GenPairX's filtering mechanism and lightweight alignment approach successfully maps 89.1% and aligns 76.1% of the reads without relying on computationally-intensive dynamic programming. We compare the proposed pipeline against the state-of-the-art software Minimap2 [79] (CPU implementation), an end-to-end GPU implementation of the BWA-MEM [80] read mapping pipeline, and the state-of-the-art ASIC-based accelerator, GenCache [68] (more details and systems are discussed in §6). Our experimental results show that GenPairX offers substantial improvements in throughput, energy efficiency, and area efficiency, achieving end-to-end throughput per Watt improvements of 1575× and 1.43× compared to state-of-the-art CPU and ASIC read mappers, respectively. GenPairX improves throughput per area by 958× and 2.38× compared to state-of-the-art CPU and ASIC mappers while maintaining or improving mapping accuracy. This paper makes the following contributions: • We perform an extensive profiling of paired-end read mapping and identify limitations and inefficiencies of commonlyused mapping tools when handling paired-end reads.

• We introduce GenPair, a new read mapping algorithm for paired-end reads. GenPair is based on a novel hash-based filter tailored for paired-end reads and a lightweight alignment algorithm that leverages the patterns in the edit variations of the expected alignments to replace unnecessary expensive dynamic programming operations with more efficient ones.

• We introduce GenPairX, the first hardware-algorithm codesigned accelerator for paired-end read mapping. GenPairX addresses the memory bottleneck inherent in GenPair by designing an optimized data structure, as well as an efficient hardware pipeline to achieve high performance and power efficiency.

• We evaluate GenPairX against two state-of-the-art software and hardware read mappers, demonstrating large performance per Watt and performance per area improvements while maintaining or improving accuracy.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines