MagiCache: A Virtual In-Cache Computing Engine
Renhao Fan, Yikai Cui, Weike Li, Mingyu Wang, Zhaolin Li
Abstract
The rise of data-parallel applications poses a significant challenge to the energy consumption of computing architectures.In-cache computation is a promising solution for achieving high parallelism and energy efficiency because it can eliminate data movement between the cache and the processor.Existing in-cache computing architectures transform a portion of cache arrays into computing arrays, with all rows of these arrays serving as computing lines.The remaining cache arrays are used as cachelines to store the data required by computing arrays or processors.However, in these array-level in-cache computing architectures, only a few computing lines in each computing array are active at runtime while the others are idle, which incurs severe cache capacity loss and space underutilization.In addition, bursty memory accesses of data-parallel applications also cause significant in-cache data movement latency.To address these problems, we propose MagiCache, a virtual in-cache computing engine.First, we design a novel cacheline-level in-cache computing architecture in which each cache array can configure some rows as computing lines and the other rows as cachelines with negligible overhead.Second, a virtual engine is further designed on this novel architecture to dynamically allocate different rows of each array as computing lines or cachelines based on runtime computation and storage requirements, thus realizing efficient cacheline-level space management.Finally, we present an instruction chaining technique to overlap the bursty access latency by enabling asynchronous execution of computing arrays.Evaluation results show that MagiCache achieves a 1.19x-1.61xspeedup over the state-of-the-art in-cache computing architectures with 6.5 KB of additional storage.Our cacheline-level space * These authors contributed equally to this work.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9d55f04b-6111-49d3-8d05-c2e84e81696eRelated papers
- Multi-Dimensional Vector ISA Extension for Mobile In-Cache ComputingAlireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu et al.HPCA 2025 · 4 citations
- Execution Sequence Optimization for Processing In-Memory using Parallel Data PreparationMuhammad Rashedul Haq Rashed, Sven Thijssen, Dominic Simon, Sumit Jha et al.DAC 2024 · 2 citations
- ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory based SSDsCongming Gao, Xin Xin, Youyou Lu, Youtao Zhang et al.MICRO 2021 · 42 citations
- ARCANE: Adaptive RISC-V Cache Architecture for Near-memory ExtensionsVincenzo Petrolo, Flavia Guella, Michele Caon, Pasquale Davide Schiavone et al.DAC 2025 · 1 citation
- ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAMXin Xin, Youtao Zhang, Jun YangHPCA 2020 · 84 citations
