CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale Systems
Sagar Karandikar, Aniruddha N. Udipi, Junsun Choi, Joonho Whangbo, Jerry Zhao, Svilen Kanev, Edwin Lim, Jyrki Alakuijala, Vrishab Madduri, Yakun Sophia Shao, Borivoje Nikolic, Krste Asanovic, Parthasarathy Ranganathan
Abstract
General-purpose lossless data compression and decompression ("(de)compression") are used widely in hyperscale systems and are key "datacenter taxes". However, designing optimal hardware compression and decompression processing units ("CDPUs") is challenging due to the variety of algorithms deployed, input data characteristics, and evolving costs of CPU cycles, network bandwidth, and memory/storage capacities.
To navigate this vast design space, we present the first largescale data-driven analysis of (de)compression usage at a major cloud provider by profiling Google's datacenter fleet. We find that (de)compression consumes 2.9% of fleet CPU cycles and 10-50% of cycles in key services. Demand is also artificially limited; 95% of bytes compressed in the fleet use less capable algorithms to reduce compute, motivating a CDPU that changes cost vs. size tradeoffs.
Prior work has improved the microarchitectural state-of-the-art for CDPUs supporting various algorithms in fixed contexts. However, we find that higher-level design parameters like CDPU placement, hash table sizing, history window sizes, and more have as significant of an impact on the viability of CDPU integration, but are not well-studied. Thus, we present the first end-to-end design/evaluation framework for CDPUs, including: 1. An open-source RTLbased CDPU generator that supports many run-time and compiletime parameters. 2. Integration into an open-source RISC-V SoC for
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06bbd8f6-7f41-4caf-a54d-9f1657b909daCited by top-tier papers9
- Sabre: Hardware-Accelerated Snapshot Compression for Serverless MicroVMsNikita Lazarev, Varun Gohil, James Tsai, Andy Anderson et al.OSDI 2024 · 13 citations
- Reviving In-Storage Hardware Compression on ZNS SSDs through Host-SSD CollaborationYingjia Wang, Tao Lu, Yuhong Liang, Xiang Chen et al.HPCA 2025 · 7 citations
- HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative ComputingJinghan Huang, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong et al.ISCA 2024 · 7 citations
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li et al.HPCA 2025 · 6 citations
- SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence AnalysisNika Mansouri-Ghiasi, Talu Güloglu, Harun Mustafa, Can Firtina et al.HPCA 2026 · 3 citations
Builds on6
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang et al.ASPLOS 2022 · 103 citations
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland et al.ASPLOS 2020 · 43 citations
- Warehouse-scale video acceleration: co-design and deployment in the wildParthasarathy Ranganathan, Daniel Stodolsky, Jeff Calow, Jeremy Dorfman et al.ASPLOS 2021 · 43 citations
- A Hardware Accelerator for Protocol BuffersSagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao et al.MICRO 2021 · 42 citations
Related papers
- ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling InsightsTao Lu, Jiapin Wang, Yelin Shan, Xiangping Zhang et al.EuroSys 2026
- cuSZp2: A GPU Lossy Compressor with Extreme Throughput and Optimized Compression RatioYafan Huang, Sheng Di, Guanpeng Li, Franck CappelloSC 2024 · 29 citations
- Efficient Lossless Compression of Scientific Floating-Point Data on CPUs and GPUsNoushin Azami, Alex Fallin, Martin BurtscherASPLOS 2025 · 18 citations
- cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with Optimized End-to-End PerformanceYafan Huang, Sheng Di, Xiaodong Yu, Guanpeng Li et al.SC 2023 · 52 citations
- Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless OrchestrationShixun Wu, Jinwen Pan, Jinyang Liu, Jiannan Tian et al.SC 2025 · 6 citations
