YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
Zihao Xuan, Yuxuan Yang, Wei Xuan, Zijia Su, Song Chen, Yi Kang
Abstract
In this paper, we further explore the potential of analog in-memory computing (AiMC) and introduce an innovative artificial intelligence (AI) accelerator architecture named YOCO, featuring three key proposals: (1) YOCO proposes a novel 8-bit in-situ multiply arithmetic (IMA) achieving 123.8 TOPS/W energy-efficiency and 34.9 TOPS throughput through efficient charge-domain computation and time-domain accumulation mechanism. (2) YOCO employs a hybrid ReRAM-SRAM memory structure to balance computational efficiency and storage density. (3) YOCO tailors an IMC-friendly attention computing flow with an efficient pipeline to accelerate the inference of transformer-based AI models. Compared to three SOTA baselines, YOCO on average improves energy efficiency by up to and throughput by up to across transformer models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0705943d-21eb-4d11-85bf-0af353788d0cBuilds on2
- Timely: Pushing Data Movements And Interfaces In Pim Accelerators Towards Local And In Time DomainWeitao Li, Pengfei Xu, Yang Zhao, Haitong Li et al.ISCA 2020 · 86 citations
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 45 citations
Related papers
- A Charge-Sharing based 8T SRAM In-Memory Computing for Edge DNN AccelerationKyeongho Lee, Sungsoo Cheon, Joongho Jo, Woong Choi et al.DAC 2021 · 34 citations
- An In-Memory Computing Accelerator with Reconfigurable Dataflow for Multi-Scale Vision Transformer with Hybrid TopologyZhiyuan Chen, Yufei Ma, Keyi Li, Yifan Jia et al.DAC 2024 · 2 citations
- Improving the Efficiency of In-Memory-Computing Macro with a Hybrid Analog-Digital Computing Mode for Lossless Neural Network InferenceQilin Zheng, Ziru Li, Jonathan Ku, Yitu Wang et al.DAC 2024 · 2 citations
- Reshape and Adapt for Output Quantization (RAOQ): Quantization-aware Training for In-memory Computing SystemsBonan Zhang, Chia-Yu Chen, Naveen VermaICML 2024 · 9 citations
- InfoX: an energy-efficient ReRAM accelerator design with information-lossless low-bit ADCsYintao He, Songyun Qu, Ying Wang, Bing Li et al.DAC 2022 · 10 citations
