A near-storage framework for boosted data preprocessing of mass spectrum clustering
Weihong Xu, Jaeyoung Kang, Tajana Rosing
摘要
Mass spectrometry (MS) has been a key to proteomics and metabolomics due to its unique ability to identify and analyze protein structures. Modern MS equipment generates massive amount of tandem mass spectra with high redundancy, making spectral analysis the major bottleneck in design of new medicines. Mass spectrum clustering is one promising solution as it greatly reduces data redundancy and boosts protein identification. However, state-of-the-art MS tools take many hours to run spectrum clustering. Spectra loading and preprocessing consumes average 82% execution time and energy during clustering. We propose a near-storage framework, MSAS, to speed up spectrum preprocessing. Instead of loading data into host memory and CPU, MSAS processes spectra near storage, thus reducing the expensive cost of data movement. We present two types of accelerators that leverage internal bandwidth at two storage levels: SSD and channel. The accelerators are optimized to match the data rate at each storage level with negligible overhead. Our results demonstrate that the channel-level design yields the best performance improvement for preprocessing - it is up to 187X and 1.8X faster than the CPU and the state-of-the-art in-storage computing solution, INSIDER, respectively. After integrating channel-level MSAS into existing MS clustering tools, we measure system level improvements in speed of 3.5X to 9.8X with 2.8X to 11.9X better energy efficiency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLMZhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai 等MICRO 2024 · 被引用 29 次
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang 等HPCA 2024 · 被引用 27 次
相关 Paper
- SpectraFlux: Harnessing the Flow of Multi-FPGA in Mass Spectrometry ClusteringTianqi Zhang, Neha Prakriya, Sumukh Pinge, Jason Cong 等DAC 2024 · 被引用 1 次
- MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage ProcessingNika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer 等ISCA 2024 · 被引用 15 次
- Sidekick: Near Data Processing for Clustering Enhanced by Automatic Memory DisaggregationSanghoon Lee, Jongho Park, Minho Ha, Byungil Koh 等DAC 2023 · 被引用 2 次
- DUAL: Acceleration of Clustering Algorithms using Digital-based Processing In-MemoryMohsen Imani, Saikishan Pampana, Saransh Gupta, Minxuan Zhou 等MICRO 2020 · 被引用 91 次
- P-Massive: A Real-Time Search Engine for a Multi-Terabyte Mass Spectrometry DatabaseNarangerelt Batsoyol, Benjamin S. Pullman, Mingxun Wang, Nuno Bandeira 等SC 2022 · 被引用 17 次
