Ares-Flash: Efficient Parallel Integer Arithmetic Operations Using NAND Flash Memory
Jian Chen, Congming Gao, Youyou Lu, Yuhao Zhang, Jiwu Shu
Abstract
In-Flash Processing (IFP) has been proposed in recent years to realize computation ability inside NAND flash memory. Distinguished from processing-in-memory (PIM) and in-storage-processing (ISP), IFP reduces the data movement starting from the most bottom flash memory medium. It is especially beneficial to those applications that require high processing parallelism (e.g., huge databases, large-scale image processing, etc.). However, IFP works often require expensive extra hardware modification, which brings lots of energy dissipation and area overhead. Recent works, Parabit and Flash-Cosmos, try to implement calculations using flash memory with minimum hardware modification. But their accomplishments still remain at the level of simple bitwise operations, and thus it prevents their work from being widely adopted. In this work, we propose Ares-Flash, a new technology using flash memory to support more complex integer arithmetic operations (e.g., addition, accumulation, and multiplication). Ares leverages the designed page buffer to perform basic mechanisms: full-adder logic and bit-shift. Moreover, we construct computational operations using dedicated control sequences in the page buffer upon two basic mechanisms. Our experimental results indicate that Ares is highly efficient and significantly mitigates data movement from storage to memory or computing units (e.g., CPUs, GPUs). Quantitatively, Ares averagely improves performance and energy efficiency by 8.58×/9.89× and 97.9×/14× compared to the out-storage-processing(OSP)/instorage-processing(ISP) under accumulation tasks with real workloads. It also improves 8.2×/4.5× and 89×/13.2× when performing vector-vector multiplication on real-world workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5bc9e2d8-87ba-4b14-a5da-520ce673d933Cited by top-tier papers3
- CIPHERMATCH: Accelerating Homomorphic Encryption-Based String Matching via Memory-Efficient Data Packing and In-Flash ProcessingMayank Kabra, Rakesh Nadig, Harshita Gupta, Rahul Bera et al.ASPLOS 2025 · 9 citations
- Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesRakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta et al.HPCA 2026 · 2 citations
- DARTH-PUM: A Hybrid Processing-Using-Memory ArchitectureRyan Wong, Ben Feinberg, Saugata GhoseASPLOS 2026 · 1 citation
Related papers
- Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash MemoryJisung Park, Roknoddin Azizi, Geraldo F. Oliveira, Mohammad Sadrosadati et al.MICRO 2022 · 53 citations
- CrossBit: Bitwise Computing in NAND Flash Memory with Inter-Bitline Data CommunicationHyunjin Kim, Seunghwan Song, Sukhyun Choi, Jeongin Choe et al.MICRO 2025 · 3 citations
- PIPF-DRAM: processing in precharge-free DRAMNezam Rohbani, Mohammad Arman Soleimani, Hamid Sarbazi-AzadDAC 2022 · 7 citations
- vPIM: Efficient Virtual Address Translation for Scalable Processing-in-Memory ArchitecturesAmel Fatima, Sihang Liu, Korakit Seemakhupt, Rachata Ausavarungnirun et al.DAC 2023 · 6 citations
- Hyper-Ap: Enhancing Associative Processing Through A Full-Stack OptimizationYue Zha, Jing LiISCA 2020 · 31 citations
