Scheduling Linux Threads under I/O Chiplet Wall Using cSwitch
Seunghyun An, Joontaek Oh, Ming Liu
Abstract
Chiplet-based processors have become the dominant architecture for modern servers, partitioning functionality across compute chiplets and a centralized I/O chiplet. While this design improves scalability, it introduces a new bottleneck— the I/O Chiplet Wall—where off-chip memory and device requests contend along a shared communication path through the I/O chiplet and interconnect fabric. Our characterization of an AMD EPYC processor shows that this bottleneck significantly impacts performance: it introduces non-trivial latency overheads, caps memory bandwidth, propagates congestion back to cores, reduces per-core effective bandwidth, and lacks traffic control. However, existing OS schedulers remain unaware of this bottleneck, leading to pathological scheduling behaviors and substantial inefficiencies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a5f793b0-d13e-4eaa-8c91-ebcb36ff9f17Related papers
- Understanding and Profiling the Accelerator Chiplet Network Using PingPointJunyeol Ryu, Ming Liu, Matthew D. SinclairSIGCOMM 2026
- COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM AcceleratorsTaishu Sheng, Guangyu Sun, Dezun DongHPCA 2026
- CHARM: Chiplet Heterogeneity-Aware Runtime Mapping SystemAlessandro Fogli, Bo Zhao, Peter R. Pietzuch, Jana GicevaEuroSys 2026 · 2 citations
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 3 citations
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu et al.DAC 2024
