Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash Memory
Jisung Park, Roknoddin Azizi, Geraldo F. Oliveira, Mohammad Sadrosadati, Rakesh Nadig, David Novo, Juan Gómez-Luna, Myungsuk Kim, Onur Mutlu
摘要
Bulk bitwise operations, i. e., bitwise operations on large bit vectors, are prevalent in a wide range of important application domains, including databases, graph processing, genome analysis, cryptography, and hyper-dimensional computing. In conventional systems, the performance and energy efficiency of bulk bitwise operations are bottlenecked by data movement between the compute units (e.g., CPUs and GPUs) and the memory hierarchy. In-flash processing (i. e., processing data inside NAND flash chips) has a high potential to accelerate bulk bitwise operations by fundamentally reducing data movement through the entire memory hierarchy, especially when the processed data does not fit into main memory. We identify two key limitations of the state-of-the-art in-flash processing technique for bulk bitwise operations; (i) it falls short of maximally exploiting the bit-level parallelism of bulk bitwise operations that could be enabled by leveraging the unique cell-array architecture and operating principles of NAND flash memory; (ii) it is unreliable because it is not designed to take into account the highly error-prone nature of NAND flash memory. We propose Flash-Cosmos (Flash C omputation with-O ne-S hot M ulti-O perand S ensing), a new in-flash processing technique that significantly increases the performance and energy efficiency of bulk bitwise operations while providing high reliability. Flash-Cosmos introduces two key mechanisms that can be easily supported in modern NAND flash chips: (i) M ulti-W ordline S ensing (MWS), which enables bulk bitwise operations on a large number of operands (tens of operands) with a single sensing operation, and (ii) E nhanced S LC-mode P rogramming (ESP), which enables reliable computation inside NAND flash memory. We demonstrate the feasibility of performing bulk bitwise operations with high reliability in Flash-Cosmos by testing 160 real 3D NAND flash chips. Our evaluation shows that Flash-Cosmos improves average performance and energy efficiency by and , respectively, over the state-of-the-art in-flash/outside-storage processing techniques across three real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup TablesJoão Dinis Ferreira, Gabriel Falcão, Juan Gómez-Luna, Mohammed Alser 等MICRO 2022 · 被引用 60 次
- Venice: Improving Solid-State Drive Parallelism at Low Cost via Conflict-Free AccessesRakesh Nadig, Mohammad Sadrosadati, Haiyu Mao, Nika Mansouri-Ghiasi 等ISCA 2023 · 被引用 29 次
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John 等ASPLOS 2023 · 被引用 20 次
- MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage ProcessingNika Mansouri-Ghiasi, Mohammad Sadrosadati, Harun Mustafa, Arvid Gollwitzer 等ISCA 2024 · 被引用 15 次
- CHOPPER: A Compiler Infrastructure for Programmable Bit-serial SIMD Processing Using Memory in DRAMXiangjun Peng, Yaohua Wang, Ming-Chang YangHPCA 2023 · 被引用 15 次
它引用的顶会 Paper17
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira 等ASPLOS 2021 · 被引用 182 次
- ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAMXin Xin, Youtao Zhang, Jun YangHPCA 2020 · 被引用 84 次
- SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory SystemsMaciej Besta, Raghavendra Kanakagiri, Grzegorz Kwasniewski, Rachata Ausavarungnirun 等MICRO 2021 · 被引用 78 次
- GenStore: a high-performance in-storage processing system for genome sequence analysisNika Mansouri-Ghiasi, Jisung Park, Harun Mustafa, Jeremie S. Kim 等ASPLOS 2022 · 被引用 74 次
- Reducing solid-state drive read latency by optimizing read-retryJisung Park, Myungsuk Kim, Myoungjun Chun, Lois Orosa 等ASPLOS 2021 · 被引用 66 次
相关 Paper
- Ares-Flash: Efficient Parallel Integer Arithmetic Operations Using NAND Flash MemoryJian Chen, Congming Gao, Youyou Lu, Yuhao Zhang 等MICRO 2024 · 被引用 7 次
- CrossBit: Bitwise Computing in NAND Flash Memory with Inter-Bitline Data CommunicationHyunjin Kim, Seunghwan Song, Sukhyun Choi, Jeongin Choe 等MICRO 2025 · 被引用 3 次
- A Compute-in-Memory Architecture Compatible with 3D NAND Flash that Parallelly Activates Multi-LayersLiang Zhao, Chu Yan, Fan Yang, Shifan Gao 等DAC 2021 · 被引用 15 次
- SHERLOCK: Scheduling Efficient and Reliable Bulk Bitwise Operations in NVMsHamid Farzaneh, João Paulo C. de Lima, Ali Nezhadi Khelejani, Asif Ali Khan 等DAC 2024 · 被引用 2 次
- WISEDRAM: A Reliable Bitwise In-DRAM AcceleratorMohammad Arman Soleimani, Nezam Rohbani, Adrián Cristal Kestelman, Osman S. Unsal 等DAC 2025 · 被引用 1 次
