CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory Ordering
Yanpeng Yu, Nicolai Oswald, Anurag Khandelwal
Abstract
Increasingly, multi-processing unit (PU) systems (e.g., CPU-GPU, multi-CPU, multi-GPU, etc.) are embracing cache-coherent shared memory to facilitate inter-PU communication. The coherence protocols in these systems support write-through accesses that place the data directly at the LLC to enable efficient producer-consumer communications pervasive in AI/ML workloads. Moreover, release consistency has emerged as the standard memory model in such systems due to its programming simplicity and ability to support high performance. In today's multi-PU systems, the source processor that issues the writes also orders them to enforce release consistency, even for write-through accesses. Unfortunately, such source ordering of write-through operations results in unnecessary communications between the source processor and the LLC directory, incurring significant performance, interconnect traffic, and energy overheads for multi-PU applications.
To eliminate such communication, we present cord 1 , a novel cache coherence protocol that orders write-through accesses directly at the cache directory. cord employs several novel mechanisms to minimize the metadata required for ordering traffic while efficiently scaling to multiple directories. Evaluations atop the gem5 simulator show that compared to source ordering, cord improves application performance by 24% and reduces traffic by 13% on average while incurring < 1% storage, area, and power overheads. Compared to hand-optimized message-passing implementations, cord observes a mere 3% performance overhead and 6% more traffic on average with a significantly simpler programming model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa8f561b-6a8b-41cd-bce7-d34932737799Cited by top-tier papers3
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min et al.SIGCOMM 2026 · 1 citation
- PIPM: Partial and Incremental Page Migration for Multi-host CXL Disaggregated Shared MemoryGangqi Huang, Heiner Litz, Yuanchao XuASPLOS 2026 · 1 citation
- HarborMaster: Rollback Detection for Trusted Distributed ComputingShubham Mishra, Alexander Thomas, Nurzhan Abdrassilov, Kaiyuan Chen et al.VLDB 2026
Builds on9
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsHuaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst et al.ASPLOS 2023 · 328 citations
- MIND: In-Network Memory Management for Disaggregated Data CentersSeung-Seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal et al.SOSP 2021 · 51 citations
- Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared MemoryMingxing Zhang, Teng Ma, Jinqi Hua, Zheng Liu et al.SOSP 2023 · 38 citations
- HMG: Extending Cache Coherence Protocols Across Modern Hierarchical Multi-GPU SystemsXiaowei Ren, Daniel Lustig, Evgeny Bolotin, Aamer Jaleel et al.HPCA 2020 · 38 citations
Related papers
- PhasedStore: Supporting High-Performance Write-Through Cache-Coherence Protocols Under TSOBurak Ocalan, Chloe Alverti, Shashwat Jaiswal, Antonis Psistakis et al.HPCA 2026
- Dorado: Clustered Hardware Cache Coherence for 1,000+ CoresJovan Stojkovic, Abraham Farrell, Gerasimos Gerogiannis, Zhangxiaowen Gong et al.ISCA 2026 · 1 citation
- XTRA: Unifying Cache Coherence and Concurrency Control for Distributed Transactions in a CXL PodZhijun Yang, Yu Hua, Ming Zhang, Menglei Chen et al.SOSP 2026
- A Simple Cache Coherence Scheme for Integrated CPU-GPU SystemsArdhi Wiratama Baskara Yudha, Reza Pulungan, Henry Hoffmann, Yan SolihinDAC 2020 · 3 citations
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman et al.ASPLOS 2026
