SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis
Nika Mansouri-Ghiasi, Talu Güloglu, Harun Mustafa, Can Firtina, Konstantina Koliogeorgi, Konstantinos Kanellopoulos, Haiyu Mao, Rakesh Nadig, Mohammad Sadrosadati, Jisung Park, Onur Mutlu
Abstract
King's College London 4 POSTECH Genome sequence analysis, which examines the DNA sequences of organisms, drives advances in many critical medical and biotechnological fields. Given its importance and the exponentially growing volumes of genomic sequence data, there are extensive efforts to accelerate genome sequence analysis. In this work, we demonstrate a major bottleneck that greatly limits and diminishes the benefits of state-of-the-art genome sequence analysis accelerators: the data preparation bottleneck, where genomic sequence data is stored in compressed form and needs to be first decompressed and formatted before an accelerator can operate on it. To mitigate this bottleneck, we propose SAGe, an algorithm-architecture co-design for highly-compressed storage and high-performance access of large-scale genomic sequence data. The key challenge is to improve data preparation performance while maintaining high compression ratios (comparable to genomic-specific compression algorithms) at low hardware cost. We address this challenge by leveraging key properties of genomic datasets to co-design (i) a lossless (de)compression algorithm, (ii) hardware that decompresses data with lightweight operations and efficient streaming accesses, (iii) storage data layout, and (iv) interface commands to access data. SAGe is highly versatile, as it supports datasets from different sequencing technologies and species. Due to its lightweight design, SAGe can be seamlessly integrated with a broad range of hardware accelerators for genome sequence analysis to mitigate their data preparation bottlenecks. Our results demonstrate that SAGe improves the average end-to-end performance and energy efficiency of two state-of-the-art genome sequence analysis accelerators by 3.0×-32.1× and 13.0×-34.0×, respectively, compared to when the accelerators rely on state-of-the-art software and hardware decompression tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af591865-407e-4de1-8bf7-c1ca96dcd1b6Builds on25
- Benchmarking Learned IndexesRyan Marcus, Andreas Kipf, Alexander van Renen, Mihail Stoian et al.VLDB 2021 · 185 citations
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim et al.ISCA 2022 · 66 citations
- Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash MemoryJisung Park, Roknoddin Azizi, Geraldo F. Oliveira, Mohammad Sadrosadati et al.MICRO 2022 · 53 citations
- SeedEx: A Genome Sequencing Accelerator for Optimal Alignments in Subminimal SpaceDaichi Fujiki, Shunhao Wu, Nathan Ozog, Kush Goliya et al.MICRO 2020 · 52 citations
- SeGraM: a universal hardware accelerator for genomic sequence-to-graph and sequence-to-sequence mappingDamla Senol Cali, Konstantinos Kanellopoulos, Joël Lindegger, Zülal Bingöl et al.ISCA 2022 · 38 citations
Related papers
- QUETZAL: Vector Acceleration Framework for Modern Genome Sequence Analysis AlgorithmsJulian Pavon, Iván Vargas Valdivieso, Carlos Rojas, César Hernández et al.ISCA 2024 · 5 citations
- Genesis: A Hardware Acceleration Framework for Genomic Data AnalysisTae Jun Ham, David Bruns-Smith, Brendan Sweeney, Yejin Lee et al.ISCA 2020 · 23 citations
- GRAINS: Enabling High-Performance and Low-Cost Graph-Based Genome Analysis via Storage-Aware Algorithm-Architecture Co-DesignNika Mansouri-Ghiasi, Harun Mustafa, Talu Güloglu, Rakesh Nadig et al.ISCA 2026 · 4 citations
- GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence AnalysisDamla Senol Cali, Gurpreet S. Kalsi, Zülal Bingöl, Can Firtina et al.MICRO 2020 · 23 citations
- NvWa: Enhancing Sequence Alignment Accelerator Throughput via Hardware SchedulingYewen Li, Xueqi Li, Ruihao Gao, Wanqi Liu et al.HPCA 2023 · 6 citations
