O-2A: Low Overhead DNN Compression with Outlier-Aware Approximation
Nguyen-Dong Ho, Minh-Son Le, Ik-Joon Chang
摘要
We present a low-latency DNN compression technique to reduce DRAM energy, significant in DNN inferences, namely Outlier-Aware Approximation (O-2A) coding. This technique compresses 8-bit integer, de-facto standard of DNN inferences, to 6-bit without degrading the accuracies of DNNs. The hardware for the O-2A coding can be easily embedded to DRAM controllers due to small overhead. In an Eyeriss platform, the O-2A coding improves both DRAM energy and system performance by 18 20%. The O-2A coding enables us to implement an error-correction scheme without additional parity overhead, opening the possibility of an approximate DRAM to simultaneously reduce DRAM accessing and refresh energy.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Effective zero compression on ReRAM-based sparse DNN acceleratorsHoon Shin, Rihae Park, Seung Yul Lee, Yeonhong Park 等DAC 2022 · 被引用 11 次
- ADROIT: An Adaptive Dynamic Refresh Optimization Framework for DRAM Energy Saving In DNN TrainingXinhan Lin, Liang Sun, Fengbin Tu, Leibo Liu 等DAC 2021 · 被引用 2 次
- DRQ: Dynamic Region-based Quantization for Deep Neural Network AccelerationZhuoran Song, Bangqi Fu, Feiyang Wu, Zhaoming Jiang 等ISCA 2020 · 被引用 92 次
- A Memory-Efficient Edge Inference Accelerator with XOR-based Model CompressionHyunseung Lee, Jihoon Hong, Soosung Kim, Seung Yul Lee 等DAC 2023 · 被引用 4 次
- Reducing Bit Writes in Non-volatile Main Memory by Similarity-aware CompressionZhangyu Chen, Yu Hua, Pengfei Zuo, Yuanyuan Sun 等DAC 2020 · 被引用 7 次
