Adaptable Register File Organization for Vector Processors
Cristóbal Ramírez Lazo, Enrico Reggiani, Carlos Rojas Morales, Roger Figueras Bagué, Luis A. Villa Vargas, Marco Antonio Ramírez Salinas, Mateo Valero Cortés, Osman Sabri Unsal, Adrián Cristal
Abstract
Contemporary Vector Processors (VPs) are de-signed either for short vector lengths, e.g., Fujitsu A64FX with 512-bit ARM SVE vector support, or long vectors, e.g., NEC Aurora Tsubasa with 16Kbits Maximum Vector Length (MVL1). Unfortunately, both approaches have drawbacks. On the one hand, short vector length VP designs struggle to provide high efficiency for applications featuring long vectors with high Data Level Parallelism (DLP). On the other hand, long vector VP designs waste resources and underutilize the Vector Register File (VRF) when executing low DLP applications with short vector lengths. Therefore, those long vector VP implementations are limited to a specialized subset of applications, where relatively high DLP must be present to achieve excellent performance with high efficiency. Modern scientific applications are getting more diverse, and the vector lengths in those applications vary widely. To overcome these limitations, we propose an Adaptable Vector Architecture (AVA) that leads to having the best of both worlds. AVA is designed for short vectors (MVL=16 elements) and is thus area and energy-efficient. However, AVA has the functionality to reconfigure the MVL, thereby allowing to exploit the benefits of having a longer vector of up to 128 elements microarchitecture when abundant DLP is present. We model AVA on the gem5 simulator and evaluate AVA performance with six applications taken from the RiVEC Benchmark Suite. To obtain area and power consumption metrics, we model AVA on McPAT for 22nm technology. Our results show that by reconfiguring our small VRF (8KB) plus our novel issue queue scheme, AVA yields a 2X speedup over the default configuration for short vectors. Additionally, AVA shows competitive performance when compared to a long vector VP, while saving 50% of area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfc6a0a0-1201-48b4-b03d-2c2564812e27Related papers
- Unlimited Vector Extension with Data Streaming SupportJoao Mario Domingos, Nuno Neves, Nuno Roma, Pedro TomásISCA 2021 · 31 citations
- big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on ChipTuan Ta, Khalid Al-Hawaj, Nick Cebry, Yanghui Ou et al.MICRO 2022 · 10 citations
- Multi-Dimensional Vector ISA Extension for Mobile In-Cache ComputingAlireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu et al.HPCA 2025 · 4 citations
- Co-design for A64FX manycore processor and "Fugaku"Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama et al.SC 2020 · 112 citations
- Improving Predication Efficiency through Compaction/Restoration of SIMD InstructionsAdrián Barredo, Juan M. Cebrian, Miquel Moretó, Marc Casas et al.HPCA 2020 · 8 citations
