Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
Alireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu, Nishil Talati, Scott A. Mahlke, Reetuparna Das
Abstract
In-cache computing technology transforms existing caches into long-vector compute units and offers low-cost alternatives to building expensive vector engines for mobile CPUs. Unfortunately, existing long-vector Instruction Set Architecture (ISA) extensions, such as RISC-V Vector Extension (RVV) and Arm Scalable Vector Extension (SVE), provide only onedimensional strided and random memory accesses. While this is sufficient for typical vector engines, it fails to effectively utilize the large Single Instruction, Multiple Data (SIMD) widths of incache vector engines. This is because mobile data-parallel kernels expose limited parallelism across a single dimension. Based on our analysis of mobile vector kernels, we introduce a long-vector Multi-dimensional Vector ISA Extension (MVE) for mobile in-cache computing. MVE achieves high SIMD resource utilization and enables flexible programming by abstracting cache geometry and data layout. The proposed ISA features multi-dimensional strided and random memory accesses and efficient dimension-level masked execution to encode parallelism across multiple dimensions. Using a wide range of data-parallel mobile workloads, we demonstrate that MVE offers significant performance and energy reduction benefits of and , on average, compared to the SIMD units of a commercial mobile processor, at an area overhead of 3.6%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7054682-f5b7-4314-a0ee-22441bad7e2aCited by top-tier papers2
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati et al.ISCA 2026 · 3 citations
- DARTH-PUM: A Hybrid Processing-Using-Memory ArchitectureRyan Wong, Ben Feinberg, Saugata GhoseASPLOS 2026 · 1 citation
Builds on9
- Co-design for A64FX manycore processor and "Fugaku"Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama et al.SC 2020 · 112 citations
- Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern MatchingElaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea Stan et al.HPCA 2020 · 43 citations
- CAPE: A Content-Addressable Processing EngineHelena Caminal, Kailin Yang, Srivatsa Srinivasa, Akshay Krishna Ramanathan et al.HPCA 2021 · 30 citations
- Accelerating database analytic query workloads using an associative processorHelena Caminal, Yannis Chronis, Tianshu Wu, Jignesh M. Patel et al.ISCA 2022 · 19 citations
- EVE: Ephemeral Vector EnginesKhalid Al-Hawaj, Tuan Ta, Nick Cebry, Shady Agwa et al.HPCA 2023 · 15 citations
Related papers
- big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on ChipTuan Ta, Khalid Al-Hawaj, Nick Cebry, Yanghui Ou et al.MICRO 2022 · 10 citations
- MagiCache: A Virtual In-Cache Computing EngineRenhao Fan, Yikai Cui, Weike Li, Mingyu Wang et al.ISCA 2025 · 3 citations
- Unlimited Vector Extension with Data Streaming SupportJoao Mario Domingos, Nuno Neves, Nuno Roma, Pedro TomásISCA 2021 · 31 citations
- Interleaved Multi-VectorizingZhuhe Fang, Beilei Zheng, Chuliang WengVLDB 2020 · 19 citations
- MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel InferenceRenhao Fan, Yikai Cui, Qilin Chen, Mingyu Wang et al.MICRO 2023 · 12 citations
