ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling Insights
Tao Lu, Jiapin Wang, Yelin Shan, Xiangping Zhang, Xiang Chen
摘要
Lossless compression imposes significant computational overhead on datacenters when performed on CPUs. Hardware compression and decompression processing units (CDPUs) can alleviate this overhead, but optimal algorithm selection, microarchitectural design, and system-level placement of CDPUs are still not well understood. We present the design of an ASIC-based in-storage CDPU and provide a comprehensive end-to-end evaluation against two leading ASIC accelerators, Intel QAT 8970 and QAT 4xxx. The evaluation spans three dominant CDPU placement regimes: peripheral, on-chip, and in-storage. Our results reveal: (i) acute sensitivity of throughput and latency to CDPU placement and interconnection, (ii) strong correlation between compression efficiency and data patterns/layouts, (iii) placement-driven divergences between microbenchmark gains and real-application speedups, (iv) discrepancies between module and system-level power efficiency, and (v) scalability and multi-tenant interference issues of various CDPUs. These findings motivate a placement-aware, cross-layer rethinking of hardware (de)compression for hyperscale storage infrastructures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 被引用 78 次
- Scalable Billion-point Approximate Nearest Neighbor Search Using SmartSSDsBing Tian, Haikun Liu, Zhuohui Duan, Xiaofei Liao 等USENIX ATC 2024 · 被引用 53 次
- Profiling Hyperscale Big Data ProcessingAbraham Gonzalez, Aasheesh Kolli, Samira Manabi Khan, Sihang Liu 等ISCA 2023 · 被引用 30 次
- CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale SystemsSagar Karandikar, Aniruddha N. Udipi, Junsun Choi, Joonho Whangbo 等ISCA 2023 · 被引用 26 次
- Closing the B+-tree vs. LSM-tree Write Amplification Gap on Modern Storage Hardware with Built-in Transparent CompressionYifan Qiao, Xubin Chen, Ning Zheng, Jiangpeng Li 等FAST 2022 · 被引用 23 次
相关 Paper
- MetaZip: a high-throughput and efficient accelerator for DEFLATERuihao Gao, Xueqi Li, Yewen Li, Xun Wang 等DAC 2022 · 被引用 9 次
- CereSZ: Enabling and Scaling Error-bounded Lossy Compression on Cerebras CS-2Shihui Song, Yafan Huang, Peng Jiang, Xiaodong Yu 等HPDC 2024 · 被引用 14 次
- Efficient Lossless Compression of Scientific Floating-Point Data on CPUs and GPUsNoushin Azami, Alex Fallin, Martin BurtscherASPLOS 2025 · 被引用 18 次
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu 等VLDB 2024 · 被引用 12 次
- QEI: Query Acceleration Can be Generic and Efficient in the CloudYifan Yuan, Yipeng Wang, Ren Wang, Rangeen Basu Roy Chowdhury 等HPCA 2021 · 被引用 4 次
