GRAINS: Enabling High-Performance and Low-Cost Graph-Based Genome Analysis via Storage-Aware Algorithm-Architecture Co-Design
Nika Mansouri-Ghiasi, Harun Mustafa, Talu Güloglu, Rakesh Nadig, Konstantina Koliogeorgi, Susana Rebolledo Ruiz, Marc Rautmann, Furkan Eris, Mohammad Sadrosadati, Jisung Park, Onur Mutlu
Abstract
Graph-based representations of genome sequences have emerged as a powerful approach for representing massive genomic databases in an expressive and efficient way. Compared to traditional, linear genome sequences, genome graphs enable more accurate and efficient genome analyses, particularly in complex, population-scale settings (e.g., public health, precision medicine, and agriculture). Despite their benefits, analysis on large-scale genome graphs incurs significant data movement overhead from the storage system due to accessing large amounts of low-reuse data. Processing data directly inside the storage device, where data originally resides, can be a fundamental solution for mitigating this overhead. However, none of the existing tools for graph-based genome analysis can be efficiently used inside the storage system due to the limited internal hardware resources in modern SSDs. At the same time, prior storage-centric systems developed for (i) traditional, linear non-graph-based genome analysis or (ii) conventional, non-genomic graph analysis are not suitable for the unique data structures and access patterns of graph-based genome analysis. We propose GRAINS, the first system for analysis with largescale genome graphs in storage. Through our detailed examination of typical analysis pipelines that operate on genome graphs, we perform storage-aware algorithm-architecture co-design to (i) make the graph-based genome analysis pipelines more storagefriendly and (ii) further improve performance, energy-efficiency, and cost via in-storage and in-flash processing. GRAINS's codesign is based on three key aspects. First, we propose a new batching technique and execution flow, based on unique features of genome graphs, that reduces the number of random accesses to graph nodes. Second, via in-flash and in-storage processing, we avoid transferring low-reuse or unused flash pages, preventing SSD channel and external I/O bandwidth waste. Third, to leverage the full parallelism of flash dies during in-flash processing, we design an effective, yet lightweight, scheduling technique, enabled by re-purposing the existing SSD structures. GRAINS's design is versatile and flexible as it supports key operations on genome graphs and can be integrated in various analysis pipelines. GRAINS provides speedup energy reduction) over the state-of-the-art software baselines, and speedup energy reduction) over a hardware-accelerated baseline.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1003da89-a6c5-4b24-b704-949811619e16Related papers
- MeG2: In-Memory Acceleration for Genome Graphs AnalysisYu Huang, Long Zheng, Haifeng Liu, Zhuoran Zhou et al.DAC 2023 · 3 citations
- MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage ProcessingNika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer et al.ISCA 2024 · 15 citations
- FlashGNN: An In-SSD Accelerator for GNN TrainingFuping Niu, Jianhui Yue, Jiangqiu Shen, Xiaofei Liao et al.HPCA 2024 · 13 citations
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang et al.HPCA 2024 · 27 citations
- SeGraM: a universal hardware accelerator for genomic sequence-to-graph and sequence-to-sequence mappingDamla Senol Cali, Konstantinos Kanellopoulos, Joël Lindegger, Zülal Bingöl et al.ISCA 2022 · 38 citations
