YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
Zihao Xuan, Yuxuan Yang, Wei Xuan, Zijia Su, Song Chen, Yi Kang
摘要
In this paper, we further explore the potential of analog in-memory computing (AiMC) and introduce an innovative artificial intelligence (AI) accelerator architecture named YOCO, featuring three key proposals: (1) YOCO proposes a novel 8-bit in-situ multiply arithmetic (IMA) achieving 123.8 TOPS/W energy-efficiency and 34.9 TOPS throughput through efficient charge-domain computation and time-domain accumulation mechanism. (2) YOCO employs a hybrid ReRAM-SRAM memory structure to balance computational efficiency and storage density. (3) YOCO tailors an IMC-friendly attention computing flow with an efficient pipeline to accelerate the inference of transformer-based AI models. Compared to three SOTA baselines, YOCO on average improves energy efficiency by up to and throughput by up to across transformer models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Timely: Pushing Data Movements And Interfaces In Pim Accelerators Towards Local And In Time DomainWeitao Li, Pengfei Xu, Yang Zhao, Haitong Li 等ISCA 2020 · 被引用 86 次
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 被引用 45 次
相关 Paper
- A Charge-Sharing based 8T SRAM In-Memory Computing for Edge DNN AccelerationKyeongho Lee, Sungsoo Cheon, Joongho Jo, Woong Choi 等DAC 2021 · 被引用 34 次
- An In-Memory Computing Accelerator with Reconfigurable Dataflow for Multi-Scale Vision Transformer with Hybrid TopologyZhiyuan Chen, Yufei Ma, Keyi Li, Yifan Jia 等DAC 2024 · 被引用 2 次
- Improving the Efficiency of In-Memory-Computing Macro with a Hybrid Analog-Digital Computing Mode for Lossless Neural Network InferenceQilin Zheng, Ziru Li, Jonathan Ku, Yitu Wang 等DAC 2024 · 被引用 2 次
- Reshape and Adapt for Output Quantization (RAOQ): Quantization-aware Training for In-memory Computing SystemsBonan Zhang, Chia-Yu Chen, Naveen VermaICML 2024 · 被引用 9 次
- InfoX: an energy-efficient ReRAM accelerator design with information-lossless low-bit ADCsYintao He, Songyun Qu, Ying Wang, Bing Li 等DAC 2022 · 被引用 10 次
