ASSASIN: Architecture Support for Stream Computing to Accelerate Computational Storage
Chen Zou, Andrew A. Chien
摘要
Computational storage adds computing to storage devices, providing potential benefits in offload, data-reduction, and lower energy. Successful computational SSD architectures should match growing flash bandwidth, which in turn requires high SSD DRAM memory bandwidth. This creates a memory wall scaling problem, resulting from SSDs' stringent power and cost constraints.
A survey of recent computational SSD research shows that many computational storage offloads are suited to stream computing. To exploit this opportunity, we propose a novel general-purpose computational SSD and core architecture, called ASSASIN (Architecture Support for Stream computing to Accelerate computatIoNal Storage). ASSASIN provides a unified set of compute engines between SSD DRAM and the flash array. This eliminates the SSD DRAM bottleneck by enabling direct computing on flash data streams. ASSASIN further employs a crossbar to achieve performance even when flash data layout is uneven and preserve independence for page layout decisions in the flash translation layer. With stream buffers and scratchpad memories, ASSASIN core's memory hierarchy and instruction set extensions provide superior low-latency access at low-power and effectively keep streaming flash data out of the in-SSD cache-DRAM memory hierarchy, thereby solving the memory wall.
Evaluation shows that ASSASIN delivers 1.5x -2.4x speedup for offloaded functions compared to state-of-the-art computational SSD architectures. Further, ASSASIN's streaming approach yields 2.0x power efficiency and 3.2x area efficiency improvement. And these performance benefits at the level of computational SSDs translate to 1.1x -1.5x end-to-end speedups on data analytics workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLMZhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai 等MICRO 2024 · 被引用 29 次
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang 等HPCA 2024 · 被引用 27 次
- MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage ProcessingNika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer 等ISCA 2024 · 被引用 15 次
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 被引用 8 次
- SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence AnalysisNika Mansouri-Ghiasi, Talu Güloglu, Harun Mustafa, Can Firtina 等HPCA 2026 · 被引用 3 次
它引用的顶会 Paper7
- ZNS: Avoiding the Block Interface Tax for Flash-based SSDsMatias Bjørling, Abutalib Aghayev, Hans Holmberg, Aravind Ramesh 等USENIX ATC 2021 · 被引用 221 次
- GenStore: a high-performance in-storage processing system for genome sequence analysisNika Mansouri-Ghiasi, Jisung Park, Harun Mustafa, Jeremie S. Kim 等ASPLOS 2022 · 被引用 74 次
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang 等HPCA 2021 · 被引用 62 次
- LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCsHuaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte 等ASPLOS 2020 · 被引用 58 次
- GLIST: Towards In-Storage Graph LearningCangyuan Li, Ying Wang, Cheng Liu, Shengwen Liang 等USENIX ATC 2021 · 被引用 53 次
相关 Paper
- SaS: SSD as SQL Database SystemJong-Hyeok Park, Soyee Choi, Gihwan Oh, Sang Won LeeVLDB 2021 · 被引用 13 次
- StreamCSD: SSD-Autonomous Stream Management via In-Storage Content LearningWenjie Li, Xiang Chen, Yelin Shan, Jiapin Wang 等DAC 2025
- Optimizing the Performance of NDP Operations by Retrieving File Semantics in StorageLin Li, Xianzhang Chen, Jiali Li, Jiapin Wang 等DAC 2023 · 被引用 6 次
- λ-IO: A Unified IO Stack for Computational StorageZhe Yang, Youyou Lu, Xiaojian Liao, Youmin Chen 等FAST 2023 · 被引用 54 次
- CEMU: Enabling Full-System Emulation of Computational Storage Beyond Hardware LimitsQiuyang Zhang, Jiapin Wang, You Zhou, Peng Xu 等ASPLOS 2026
