O-2A: Low Overhead DNN Compression with Outlier-Aware Approximation
Nguyen-Dong Ho, Minh-Son Le, Ik-Joon Chang
Abstract
We present a low-latency DNN compression technique to reduce DRAM energy, significant in DNN inferences, namely Outlier-Aware Approximation (O-2A) coding. This technique compresses 8-bit integer, de-facto standard of DNN inferences, to 6-bit without degrading the accuracies of DNNs. The hardware for the O-2A coding can be easily embedded to DRAM controllers due to small overhead. In an Eyeriss platform, the O-2A coding improves both DRAM energy and system performance by 18 20%. The O-2A coding enables us to implement an error-correction scheme without additional parity overhead, opening the possibility of an approximate DRAM to simultaneously reduce DRAM accessing and refresh energy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c743e17a-5c1a-46bb-8de8-9fa4d3928869Related papers
- Effective zero compression on ReRAM-based sparse DNN acceleratorsHoon Shin, Rihae Park, Seung Yul Lee, Yeonhong Park et al.DAC 2022 · 11 citations
- ADROIT: An Adaptive Dynamic Refresh Optimization Framework for DRAM Energy Saving In DNN TrainingXinhan Lin, Liang Sun, Fengbin Tu, Leibo Liu et al.DAC 2021 · 2 citations
- DRQ: Dynamic Region-based Quantization for Deep Neural Network AccelerationZhuoran Song, Bangqi Fu, Feiyang Wu, Zhaoming Jiang et al.ISCA 2020 · 92 citations
- A Memory-Efficient Edge Inference Accelerator with XOR-based Model CompressionHyunseung Lee, Jihoon Hong, Soosung Kim, Seung Yul Lee et al.DAC 2023 · 4 citations
- Reducing Bit Writes in Non-volatile Main Memory by Similarity-aware CompressionZhangyu Chen, Yu Hua, Pengfei Zuo, Yuanyuan Sun et al.DAC 2020 · 7 citations
