Profiling Hyperscale Big Data Processing
Abraham Gonzalez, Aasheesh Kolli, Samira Manabi Khan, Sihang Liu, Vidushi Dadu, Sagar Karandikar, Jichuan Chang, Krste Asanovic, Parthasarathy Ranganathan
Abstract
Computing demand continues to grow exponentially, largely driven by łbig dataž processing on hyperscale data stores. At the same time, the slowdown in Moore's law is leading the industry to embrace custom computing in large-scale systems. Taken together, these trends motivate the need to characterize live production traffic on these large data processing platforms and understand the opportunity of acceleration at scale.
This paper addresses this key need. We characterize three important production distributed database and data analytics platforms at Google to identify key hardware acceleration opportunities and perform a comprehensive limits study to understand the trade-offs among various hardware acceleration strategies.
We observe that hyperscale data processing platforms spend significant time on distributed storage and other remote work across distributed workers. Therefore, optimizing storage and remote work in addition to compute acceleration is critical for these platforms. We present a detailed breakdown of the compute-intensive functions in these platforms and identify dominant key data operations related to datacenter and systems taxes. We observe that no single accelerator can provide a significant benefit but collectively, a sea of accelerators, can accelerate many of these smaller platformspecific functions. We demonstrate the potential gains of the sea of accelerators proposal in a limits study and analytical model. We perform a comprehensive study to understand the trade-offs between accelerator location (on-chip/off-chip) and invocation model (synchronous/asynchronous). We propose and evaluate a chained accelerator execution model where identified compute-intensive functions are accelerated and pipelined to avoid invocation from * Work done while at Google.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- Sabre: Hardware-Accelerated Snapshot Compression for Serverless MicroVMsNikita Lazarev, Varun Gohil, James Tsai, Andy Anderson et al.OSDI 2024 · 13 citations
- Data Motion Acceleration: Chaining Cross-Domain Multi AcceleratorsShu-Ting Wang, Hanyang Xu, Amin Mamandipoor, Rohan Mahapatra et al.HPCA 2024 · 10 citations
- Memento: Architectural Support for Ephemeral Memory Management in Serverless EnvironmentsZiqi Wang, Kaiyang Zhao, Pei Li, Andrew Jacob et al.MICRO 2023 · 7 citations
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal PerspectiveSeokjin Go, Joongun Park, Spandan More, Hanjiang Wu et al.MICRO 2025 · 7 citations
Builds on11
- Quantifying TPC-H Choke Points and Their OptimizationsMarkus Dreseler, Martin Boissier, Tilmann Rabl, Matthias UflackerVLDB 2020 · 91 citations
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu et al.VLDB 2021 · 67 citations
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland et al.ASPLOS 2020 · 43 citations
- Warehouse-scale video acceleration: co-design and deployment in the wildParthasarathy Ranganathan, Daniel Stodolsky, Jeff Calow, Jeremy Dorfman et al.ASPLOS 2021 · 43 citations
Related papers
- Terabyte-Scale Analytics in the Blink of an EyeBowen Wu, Wei Cui, Carlo Curino, Matteo Interlandi et al.VLDB 2026 · 10 citations
- dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data ProcessingJiasheng Hu, Kaiwen Zheng, Anna Li, Sidharth Sankhe et al.VLDB 2026
- A Hardware Accelerator for Protocol BuffersSagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao et al.MICRO 2021 · 42 citations
- The Art of Efficient In-memory Query Processing on NUMA Systems: a Systematic ApproachPuya Memarzia, Suprio Ray, Virendra C. BhavsarICDE 2020 · 7 citations
- HPAC-Offload: Accelerating HPC Applications with Portable Approximate Computing on the GPUZane Fink, Konstantinos Parasyris, Giorgis Georgakoudis, Harshitha MenonSC 2023 · 3 citations
