SIMR: Single Instruction Multiple Request Processing for Energy-Efficient Data Center Microservices
Mahmoud Khairy, Ahmad Alawneh, Aaron Barnes, Timothy G. Rogers
摘要
Contemporary data center servers process thousands of similar, independent requests per minute. In the interest of programmer productivity and ease of scaling, workloads in data centers have shifted from single monolithic processes toward a micro and nanoservice software architecture. As a result, single servers are now packed with many threads executing the same, relatively small task on different data.
State-of-the-art data centers run these microservices on multi-core CPUs. However, the flexibility offered by traditional CPUs comes at an energy-efficiency cost. The Multiple Instruction Multiple Data execution model misses opportunities to aggregate the similarity in contemporary microservices. We observe that the Single Instruction Multiple Thread execution model, employed by GPUs, provides better thread scaling and has the potential to reduce frontend and memory system energy consumption. However, contemporary GPUs are ill-suited for the latency-sensitive microservice space.
To exploit the similarity in contemporary microservices, while maintaining acceptable latency, we propose the Request Processing Unit (RPU). The RPU combines elements of outof-order CPUs with lockstep thread aggregation mechanisms found in GPUs to execute microservices in a Single Instruction Multiple Request (SIMR) fashion. To complement the RPU, we also propose a SIMR-aware software stack that uses novel mechanisms to batch requests based on their predicted controlflow, split batches based on predicted latency divergence and map per-request memory allocations to maximize coalescing opportunities. Our resulting RPU system processes 5.7× more requests/joule than multi-core CPUs, while increasing single thread latency by only 1.44×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MXFaaS: Resource Sharing in Serverless Environments for Parallelism and EfficiencyJovan Stojkovic, Tianyin Xu, Hubertus Franke, Josep TorrellasISCA 2023 · 被引用 39 次
- μManycore: A Cloud-Native CPU for Tail at ScaleJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2023 · 被引用 16 次
- ThreadFuser: A SIMT Analysis Framework for MIMD ProgramsAhmad Alawneh, Ni Kang, Mahmoud Khairy, Timothy G. RogersMICRO 2024
它引用的顶会 Paper11
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 被引用 83 次
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 被引用 78 次
- Vortex: Extending the RISC-V ISA for GPGPU and 3D-GraphicsBlaise Tine, Krishna Praveen Yalamarthy, Fares Elsabbagh, Hyesoon KimMICRO 2021 · 被引用 61 次
- APOLLO: An Automated Power Modeling Framework for Runtime Power Introspection in High-Volume Commercial MicroprocessorsZhiyao Xie, Xiaoqing Xu, Matt Walker, Joshua Knebel 等MICRO 2021 · 被引用 55 次
相关 Paper
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- RPU - A Reasoning Processing UnitMatthew Joseph Adiletta, Gu-Yeon Wei, David BrooksHPCA 2026 · 被引用 2 次
- The Memory Processing Unit: A Generalized Interface for End-to-End In-Memory ExecutionMinh S. Q. Truong, Yiqiu Sun, Dawei Xiong, Amol Shah 等HPCA 2026 · 被引用 1 次
- Deadline-Aware Offloading for High-Throughput AcceleratorsTsung Tai Yeh, Matthew D. Sinclair, Bradford M. Beckmann, Timothy G. RogersHPCA 2021 · 被引用 16 次
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
