CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory Ordering
Yanpeng Yu, Nicolai Oswald, Anurag Khandelwal
摘要
Increasingly, multi-processing unit (PU) systems (e.g., CPU-GPU, multi-CPU, multi-GPU, etc.) are embracing cache-coherent shared memory to facilitate inter-PU communication. The coherence protocols in these systems support write-through accesses that place the data directly at the LLC to enable efficient producer-consumer communications pervasive in AI/ML workloads. Moreover, release consistency has emerged as the standard memory model in such systems due to its programming simplicity and ability to support high performance. In today's multi-PU systems, the source processor that issues the writes also orders them to enforce release consistency, even for write-through accesses. Unfortunately, such source ordering of write-through operations results in unnecessary communications between the source processor and the LLC directory, incurring significant performance, interconnect traffic, and energy overheads for multi-PU applications.
To eliminate such communication, we present cord 1 , a novel cache coherence protocol that orders write-through accesses directly at the cache directory. cord employs several novel mechanisms to minimize the metadata required for ordering traffic while efficiently scaling to multiple directories. Evaluations atop the gem5 simulator show that compared to source ordering, cord improves application performance by 24% and reduces traffic by 13% on average while incurring < 1% storage, area, and power overheads. Compared to hand-optimized message-passing implementations, cord observes a mere 3% performance overhead and 6% more traffic on average with a significantly simpler programming model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min 等SIGCOMM 2026 · 被引用 1 次
- PIPM: Partial and Incremental Page Migration for Multi-host CXL Disaggregated Shared MemoryGangqi Huang, Heiner Litz, Yuanchao XuASPLOS 2026 · 被引用 1 次
- HarborMaster: Rollback Detection for Trusted Distributed ComputingShubham Mishra, Alexander Thomas, Nurzhan Abdrassilov, Kaiyuan Chen 等VLDB 2026
它引用的顶会 Paper9
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst 等ASPLOS 2023 · 被引用 328 次
- MIND: In-Network Memory Management for Disaggregated Data CentersSeung-Seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal 等SOSP 2021 · 被引用 51 次
- Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared MemoryMingxing Zhang, Teng Ma, Jinqi Hua, Zheng Liu 等SOSP 2023 · 被引用 38 次
- HMG: Extending Cache Coherence Protocols Across Modern Hierarchical Multi-GPU SystemsXiaowei Ren, Daniel Lustig, Evgeny Bolotin, Aamer Jaleel 等HPCA 2020 · 被引用 38 次
相关 Paper
- PhasedStore: Supporting High-Performance Write-Through Cache-Coherence Protocols Under TSOBurak Ocalan, Chloe Alverti, Shashwat Jaiswal, Antonis Psistakis 等HPCA 2026
- Dorado: Clustered Hardware Cache Coherence for 1,000+ CoresJovan Stojkovic, Abraham Farrell, Gerasimos Gerogiannis, Zhangxiaowen Gong 等ISCA 2026 · 被引用 1 次
- XTRA: Unifying Cache Coherence and Concurrency Control for Distributed Transactions in a CXL PodZhijun Yang, Yu Hua, Ming Zhang, Menglei Chen 等SOSP 2026
- A Simple Cache Coherence Scheme for Integrated CPU-GPU SystemsArdhi Wiratama Baskara Yudha, Reza Pulungan, Henry Hoffmann, Yan SolihinDAC 2020 · 被引用 3 次
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman 等ASPLOS 2026
