A near-storage framework for boosted data preprocessing of mass spectrum clustering
Weihong Xu, Jaeyoung Kang, Tajana Rosing
Abstract
Mass spectrometry (MS) has been a key to proteomics and metabolomics due to its unique ability to identify and analyze protein structures. Modern MS equipment generates massive amount of tandem mass spectra with high redundancy, making spectral analysis the major bottleneck in design of new medicines. Mass spectrum clustering is one promising solution as it greatly reduces data redundancy and boosts protein identification. However, state-of-the-art MS tools take many hours to run spectrum clustering. Spectra loading and preprocessing consumes average 82% execution time and energy during clustering. We propose a near-storage framework, MSAS, to speed up spectrum preprocessing. Instead of loading data into host memory and CPU, MSAS processes spectra near storage, thus reducing the expensive cost of data movement. We present two types of accelerators that leverage internal bandwidth at two storage levels: SSD and channel. The accelerators are optimized to match the data rate at each storage level with negligible overhead. Our results demonstrate that the channel-level design yields the best performance improvement for preprocessing - it is up to 187X and 1.8X faster than the CPU and the state-of-the-art in-storage computing solution, INSIDER, respectively. After integrating channel-level MSAS into existing MS clustering tools, we measure system level improvements in speed of 3.5X to 9.8X with 2.8X to 11.9X better energy efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 57c8d631-cfe8-4a97-b17d-6d99fcdb9210Cited by top-tier papers2
- Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLMZhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai et al.MICRO 2024 · 29 citations
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang et al.HPCA 2024 · 27 citations
Related papers
- SpectraFlux: Harnessing the Flow of Multi-FPGA in Mass Spectrometry ClusteringTianqi Zhang, Neha Prakriya, Sumukh Pinge, Jason Cong et al.DAC 2024 · 1 citation
- MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage ProcessingNika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer et al.ISCA 2024 · 15 citations
- Sidekick: Near Data Processing for Clustering Enhanced by Automatic Memory DisaggregationSanghoon Lee, Jongho Park, Minho Ha, Byungil Koh et al.DAC 2023 · 2 citations
- DUAL: Acceleration of Clustering Algorithms using Digital-based Processing In-MemoryMohsen Imani, Saikishan Pampana, Saransh Gupta, Minxuan Zhou et al.MICRO 2020 · 91 citations
- P-Massive: A Real-Time Search Engine for a Multi-Terabyte Mass Spectrometry DatabaseNarangerelt Batsoyol, Benjamin S. Pullman, Mingxun Wang, Nuno Bandeira et al.SC 2022 · 17 citations
