PhaseWeave: Phase-Aware Execution on Heterogeneous Chiplet Architectures for Datacenters
Joshua Kim, Chaojie Zhang, Íñigo Goiri, Christopher J. Rossbach, Jovan Stojkovic
Abstract
Modern datacenter applications execute in clearly distinct phases, driven by modular microservice architectures, auxiliary “datacenter tax” operations, and diverse execution blocks within the service logic. Our characterization of datacenter workloads reveals that these millisecond-scale phases exhibit recurring patterns and have varying sensitivities to hardware resources such as core frequency, network bandwidth, and memory bandwidth. This fine-grained intra-application heterogeneity leads to inefficiencies when running on traditional homogeneous server architectures, which cannot adapt to dynamic and phasespecific resource demands. As a result, datacenter operators face a trade-off between overprovisioning resources and suffering performance bottlenecks during critical phases. We propose PhaseWeave, a heterogeneous multiple-chiplet server architecture where different chiplet types are optimized for different workload phases (e.g., compute-, memory-, or network-phases). PhaseWeave transparently predicts changes in program phases using hardware counters and the distribution of system calls. Then, it uses OS signals to migrate workloads to chiplets that are best suited for the next phase. By dynamically and transparently steering execution across specialized chiplets, PhaseWeave improves resource utilization and performance without requiring changes to user code. Full-system simulations with a diverse set of datacenter workloads show that PhaseWeave reduces tail latency of datacenter applications by 65% at high loads, increases throughput by , and improves Performance/Watt by compared to homogeneous iso-area baseline.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c1cf9677-81bf-4840-a855-9c62d7fbbca9Related papers
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 3 citations
- μManycore: A Cloud-Native CPU for Tail at ScaleJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2023 · 16 citations
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias et al.SOSP 2021 · 39 citations
- The Fast and The Frugal: Tail Latency Aware Provisioning for Coping with Load VariationsAdithya Kumar, Iyswarya Narayanan, Timothy Zhu, Anand SivasubramaniamWWW 2020 · 20 citations
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
