MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing
Nika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer, Can Firtina, Julien Eudine, Haiyu Mao, Joël Lindegger, Meryem Banu Cavlak, Mohammed Alser, Jisung Park, Onur Mutlu
Abstract
Metagenomics, the study of the genome sequences of diverse organisms in a common environment, has led to significant advances in many fields. Since the species present in a metagenomic sample are not known in advance, metagenomic analysis commonly involves the key tasks of determining the species present in a sample and their relative abundances. These tasks require searching large metagenomic databases containing information on different species’ genomes. Metagenomic analysis suffers from significant data movement overhead due to moving large amounts of low-reuse data from the storage system to the rest of the system. In-storage processing can be a fundamental solution for reducing this overhead. However, designing an in-storage processing system for metagenomics is challenging because existing approaches to metagenomic analysis cannot be directly implemented in storage effectively due to the hardware limitations of modern SSDs.We propose MegIS, the first in-storage processing system designed to significantly reduce the data movement overhead of the end-to-end metagenomic analysis pipeline. MegIS is enabled by our lightweight design that effectively leverages and orchestrates processing inside and outside the storage system. Through our detailed analysis of the end-to-end metagenomic analysis pipeline and careful hardware/software co-design, we address in-storage processing challenges for metagenomics via specialized and efficient 1) task partitioning, 2) data/computation flow coordination, 3) storage technology-aware algorithmic optimizations, 4) data mapping, and 5) lightweight in-storage accelerators. MegIS’s design is flexible, capable of supporting different types of metagenomic input datasets, and can be integrated into various metagenomic analysis pipelines. Our evaluation shows that MegIS outperforms the state-of-the-art performance- and accuracy-optimized software metagenomic tools by 2.7× – 37.2× and 6.9×–100.2×, respectively, while matching the accuracy of the accuracy-optimized tool. MegIS achieves 1.5×–5.1× speedup compared to the state-of-the-art metagenomic hardware-accelerated (using processing-in-memory) tool, while achieving significantly higher accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1224d79d-ac52-4b63-8ee2-13e3cc9254f0Cited by top-tier papers8
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage ProcessingKangqi Chen, Rakesh Nadig, Manos Frouzakis, Nika Mansouri-Ghiasi et al.ISCA 2025 · 14 citations
- CIPHERMATCH: Accelerating Homomorphic Encryption-Based String Matching via Memory-Efficient Data Packing and In-Flash ProcessingMayank Kabra, Rakesh Nadig, Harshita Gupta, Rahul Bera et al.ASPLOS 2025 · 9 citations
- SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence AnalysisNika Mansouri-Ghiasi, Talu Güloglu, Harun Mustafa, Can Firtina et al.HPCA 2026 · 3 citations
- NMP-PaK: Near-Memory Processing Acceleration of Scalable De Novo Genome AssemblyHeewoo Kim, Sanjay Sri Vallabh Singapuram, Haojie Ye, Joseph Izraelevitz et al.ISCA 2025 · 2 citations
- Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesRakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta et al.HPCA 2026 · 2 citations
Builds on25
- Benchmarking Learned IndexesRyan Marcus, Andreas Kipf, Alexander van Renen, Mihail Stoian et al.VLDB 2021 · 185 citations
- BioHD: an efficient genome sequence search platform using HyperDimensional memorizationZhuowen Zou, Hanning Chen, Prathyush Poduval, Yeseong Kim et al.ISCA 2022 · 66 citations
- SquiggleFilter: An Accelerator for Portable Virus DetectionTimothy Dunn, Harisankar Sadasivan, Jack Wadden, Kush Goliya et al.MICRO 2021 · 61 citations
- SmartSAGE: training large-scale graph neural networks using in-storage processing architecturesYunjae Lee, Jinha Chung, Minsoo RhuISCA 2022 · 57 citations
- Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash MemoryJisung Park, Roknoddin Azizi, Geraldo F. Oliveira, Mohammad Sadrosadati et al.MICRO 2022 · 53 citations
Related papers
- GRAINS: Enabling High-Performance and Low-Cost Graph-Based Genome Analysis via Storage-Aware Algorithm-Architecture Co-DesignNika Mansouri-Ghiasi, Harun Mustafa, Talu Güloglu, Rakesh Nadig et al.ISCA 2026 · 4 citations
- MeG2: In-Memory Acceleration for Genome Graphs AnalysisYu Huang, Long Zheng, Haifeng Liu, Zhuoran Zhou et al.DAC 2023 · 3 citations
- A near-storage framework for boosted data preprocessing of mass spectrum clusteringWeihong Xu, Jaeyoung Kang, Tajana RosingDAC 2022 · 10 citations
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang et al.HPCA 2024 · 27 citations
- DockerSSD: Containerized In-Storage Processing and Hardware Acceleration for Computational SSDsDonghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Myoungsoo JungHPCA 2024 · 7 citations
