Dorado: Clustered Hardware Cache Coherence for 1,000+ Cores
Jovan Stojkovic, Abraham Farrell, Gerasimos Gerogiannis, Zhangxiaowen Gong, Christopher J. Hughes, Josep Torrellas
摘要
As processors continue to grow in size, they will soon include over one thousand cores and, in at least some markets, require hardware cache coherence over all of the cores. In these systems, the costs of coherence transaction latency/traffic and directory storage will escalate. An intuitive way to contain these costs is to group cores into clusters and exploit intra-cluster locality. However, latency/traffic gains are thwarted by the need to access home directories in remote clusters, and storage reductions are limited by having to track many sharers. To address these obstacles, this paper introduces Dorado, a new directory-based coherence protocol for 1,000+ cores that exploits clusters. Dorado makes three contributions. First, while each line has a Global home directory slice, it can also have Temporary home directory slices in each of the clusters where and while it is referenced. This minimizes high-latency/traffic transactions. Second, a directory can contain different types of entries and sharer pointers, each behaving differently. To use space efficiently, Dorado allows them all to dynamically share the same hardware structures-adapting their relative space to the workload sharing patterns. Third, to support many-sharer lines with modest directory storage, Dorado introduces a simple mechanism for directory entries to grow into a shared area. Simulations of 1024 cores running a variety of workloads show that Dorado is effective. It attains an average speedup of 1.36× over a same-area limited-pointer protocol by reducing the average load latency by 46.1%. Further, Dorado stays within 1% of the performance of a full bit vector protocol while using less directory storage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- Kite: A Family of Heterogeneous Interposer Topologies Enabled via Accurate Interconnect ModelingSrikant Bharadwaj, Jieming Yin, Bradford M. Beckmann, Tushar KrishnaDAC 2020 · 被引用 94 次
- CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale SystemsSagar Karandikar, Aniruddha N. Udipi, Junsun Choi, Joonho Whangbo 等ISCA 2023 · 被引用 26 次
- Dvé: Improving DRAM Reliability and Performance On-Demand via Coherent ReplicationAdarsh Patil, Vijay Nagarajan, Rajeev Balasubramonian, Nicolai OswaldISCA 2021 · 被引用 9 次
- Zero Directory Eviction Victim: Unbounded Coherence Directory and Core Cache IsolationMainak ChaudhuriHPCA 2021 · 被引用 8 次
相关 Paper
- CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory OrderingYanpeng Yu, Nicolai Oswald, Anurag KhandelwalISCA 2025 · 被引用 3 次
- WiDir: A Wireless-Enabled Directory Cache Coherence ProtocolAntonio Franques, Apostolos Kokolis, Sergi Abadal, Vimuth Fernando 等HPCA 2021 · 被引用 11 次
- XTRA: Unifying Cache Coherence and Concurrency Control for Distributed Transactions in a CXL PodZhijun Yang, Yu Hua, Ming Zhang, Menglei Chen 等SOSP 2026
- Running Consistent Applications Closer to Users with Radical for Lower LatencyNicolaas Kaashoek, Oleg Aleksandrovich Golev, Austin T. Li, Amit Levy 等SOSP 2025
- Dorado: Scaling SmartNIC Session Tables on Commodity DDRsHeng Yu, Kai Ren, Jiajun Liang, Baozeng Zhang 等SIGCOMM 2026
