Lune

DAC2025顶会

High Energy-efficiency and Low latency In-Memory Computing using Analog Accumulator and In-Memory ADC with shared References

Junyi Yang, Shuai Dong, Zhengnan Fu, Hongyang Shang, Arindam Basu

2025年份
3被引次数

摘要

This article proposes a 256×128256 \times 128 in-memory computing array using reconfigurable in-memory analog-to-digital conversion with shared references for high area efficiency (area overhead of 3% is 9X9 X better than traditional). A dual-8T SRAM bitcell is used to achieve read-write decoupling and store ternary weights. Read World Line Under Drive enabled Cascode helps to minimize current variations producing high linearity. Multi-bit input is handled with low latency and high energyefficiency by using bit-slicing (BS) with near-memory chargesharing based binary weighted accumulator (CHA). Using noise resilient training, we show software comparable performance for a MLP on MNIST, VGG-8 on CIFAR-10, and graph attention network on Cora, with respective accuracy reductions of only 0.1%,0.8%0.1 \%, 0.8 \% and 0.5% due to non-idealities. The proposed macro demonstrates high energy/area efficiency (1146 TOPS/W, 27 TOPS /mm2/ \mathbf{m m}^{\mathbf{2}} at 1/2/1b\mathbf{1} \boldsymbol{/} \mathbf{2} \boldsymbol{/} \mathbf{1 b}) in 65 nm\mathbf{6 5 ~ n m} CMOS. It increases throughput (by 1.9X1.9 X) and linearity (by 23X23 X) compared to input pulse-width modulation by using BS and CHA. Compared to conventional BS with digital accumulation after ADC, this method has 1.7X/6.6X1.7 X / 6.6 X better energy-efficiency/throughput by reducing ADC operations.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖