TALCO: Tiling Genome Sequence Alignment Using Convergence of Traceback Pointers
Sumit Walia, Cheng Ye, Arkid Bera, Dhruvi Lodhavia, Yatish Turakhia
摘要
Pairwise sequence alignment is one of the most fundamental and computationally intensive steps in genome analysis. With the improving costs and throughput of third-generation sequencing technologies and the growing availability of whole-genome datasets, longer alignments are becoming more common in the field of bioinformatics. However, the high memory demands of long alignments create significant obstacles to hardware acceleration. Banding techniques allow recovering high-quality alignments with lower memory, but they also require more memory for long alignments than what is typically available on-chip in hardware accelerators. Recently, tiling-based hardware accelerators have made remarkable strides in accelerating sequence alignment, achieving three to four orders of magnitude improvement in alignment throughput over software tools without any restrictions on alignment length. However, it is crucial to note that existing tiling heuristics can cause the alignment quality to degrade, which is a critical concern for the wider adoption of accelerators in the field of bioinformatics. To address this issue, this paper describes TALCO - a novel method for tiling long sequence alignments, that, similar to prior tiling techniques, maintains a constant memory footprint during the acceleration step independent of alignment length. However, unlike previous techniques, TALCO also ensures optimal alignments under banding constraints. TALCO does this by leveraging the convergence of traceback paths beyond a tile to a single point on the boundary of that tile - a strategy that generalizes well to a broad set of sequence alignment algorithms. We demonstrate the advantages of TALCO by applying it to two different and widely-used banded sequence alignment algorithms, X-Drop and WFA-Adapt. To the best of our knowledge, this is the first time that a tiling technique is being applied to a non-classical algorithm for sequence alignment, such as WFA-Adapt. The TALCO tiling strategy is beneficial to both software and hardware. When implemented in software, the TALCO strategy reduces the memory requirements for X-Drop and WFA-Adapt algorithms by up to 39 × and 57 ×, respectively, and when implemented as ASIC accelerator, it provides up to 1,900 × and 2,000 × improvement in alignment throughput/watt over CPU baselines implementing the same algorithms. Compared to state-of-the-art GPU and ASIC baselines implementing tiling heuristics, TALCO provides up to 50 × and 1.1 × improvement in alignment throughput, respectively, while also maintaining a higher alignment quality. Code availability: https://github.com/TurakhiaLab/TALCO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SMX: Heterogeneous Architecture for Universal Sequence Alignment AccelerationMax Doblas, Po Jui Shih, Oscar Lostes-Cazorla, Miquel Moretó 等MICRO 2025 · 被引用 7 次
- GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read MappingJulien Eudine, Chu Li, Zhuo Cheng, Renzo Andri 等HPCA 2026 · 被引用 2 次
- DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in BioinformaticsAnshu Gupta, Yingqi Cao, Jason Liang, Yatish TurakhiaHPCA 2026
它引用的顶会 Paper9
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim 等ISCA 2022 · 被引用 66 次
- SeedEx: A Genome Sequencing Accelerator for Optimal Alignments in Subminimal SpaceDaichi Fujiki, Shunhao Wu, Nathan Ozog, Kush Goliya 等MICRO 2020 · 被引用 52 次
- SeGraM: a universal hardware accelerator for genomic sequence-to-graph and sequence-to-sequence mappingDamla Senol Cali, Konstantinos Kanellopoulos, Joël Lindegger, Zülal Bingöl 等ISCA 2022 · 被引用 38 次
- EDAM: edit distance tolerant approximate matching content addressable memoryRobert Hanhan, Esteban Garzón, Zuher Jahshan, Adam Teman 等ISCA 2022 · 被引用 37 次
- GenPIP: In-Memory Acceleration of Genome Analysis via Tight Integration of Basecalling and Read MappingHaiyu Mao, Mohammed Alser, Mohammad Sadrosadati, Can Firtina 等MICRO 2022 · 被引用 32 次
相关 Paper
- Space Efficient Sequence Alignment for SRAM-Based Computing: X-Drop on the Graphcore IPULuk Burchard, Max Xiaohang Zhao, Johannes Langguth, Aydin Buluç 等SC 2023 · 被引用 9 次
- NvWa: Enhancing Sequence Alignment Accelerator Throughput via Hardware SchedulingYewen Li, Xueqi Li, Ruihao Gao, Wanqi Liu 等HPCA 2023 · 被引用 6 次
- AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read MappingSeongyeon Park, Junguk Hong, Jaeyong Song, Hajin Kim 等PPoPP 2024 · 被引用 7 次
- Lembas: Cost-Efficient Genome Alignment with External Memory and FPGA AccelerationSeongyoung Kang, Se-Min Lim, Sang-Woo JunISCA 2026
- GMX: Instruction Set Extensions for Fast, Scalable, and Efficient Genome Sequence AlignmentMax Doblas, Oscar Lostes-Cazorla, Quim Aguado-Puig, Nick Cebry 等MICRO 2023 · 被引用 11 次
