Dorado: Clustered Hardware Cache Coherence for 1,000+ Cores
Jovan Stojkovic, Abraham Farrell, Gerasimos Gerogiannis, Zhangxiaowen Gong, Christopher J. Hughes, Josep Torrellas
Abstract
As processors continue to grow in size, they will soon include over one thousand cores and, in at least some markets, require hardware cache coherence over all of the cores. In these systems, the costs of coherence transaction latency/traffic and directory storage will escalate. An intuitive way to contain these costs is to group cores into clusters and exploit intra-cluster locality. However, latency/traffic gains are thwarted by the need to access home directories in remote clusters, and storage reductions are limited by having to track many sharers. To address these obstacles, this paper introduces Dorado, a new directory-based coherence protocol for 1,000+ cores that exploits clusters. Dorado makes three contributions. First, while each line has a Global home directory slice, it can also have Temporary home directory slices in each of the clusters where and while it is referenced. This minimizes high-latency/traffic transactions. Second, a directory can contain different types of entries and sharer pointers, each behaving differently. To use space efficiently, Dorado allows them all to dynamically share the same hardware structures-adapting their relative space to the workload sharing patterns. Third, to support many-sharer lines with modest directory storage, Dorado introduces a simple mechanism for directory entries to grow into a shared area. Simulations of 1024 cores running a variety of workloads show that Dorado is effective. It attains an average speedup of 1.36× over a same-area limited-pointer protocol by reducing the average load latency by 46.1%. Further, Dorado stays within 1% of the performance of a full bit vector protocol while using less directory storage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d07acbd7-1763-4d30-b46f-d7863d4e21f1Builds on6
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- Kite: A Family of Heterogeneous Interposer Topologies Enabled via Accurate Interconnect ModelingSrikant Bharadwaj, Jieming Yin, Bradford M. Beckmann, Tushar KrishnaDAC 2020 · 94 citations
- CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale SystemsSagar Karandikar, Aniruddha N. Udipi, Junsun Choi, Joonho Whangbo et al.ISCA 2023 · 26 citations
- Dvé: Improving DRAM Reliability and Performance On-Demand via Coherent ReplicationAdarsh Patil, Vijay Nagarajan, Rajeev Balasubramonian, Nicolai OswaldISCA 2021 · 9 citations
- Zero Directory Eviction Victim: Unbounded Coherence Directory and Core Cache IsolationMainak ChaudhuriHPCA 2021 · 8 citations
Related papers
- CORD: Low-Latency, Bandwidth-Efficient and Scalable Release Consistency via Directory OrderingYanpeng Yu, Nicolai Oswald, Anurag KhandelwalISCA 2025 · 3 citations
- WiDir: A Wireless-Enabled Directory Cache Coherence ProtocolAntonio Franques, Apostolos Kokolis, Sergi Abadal, Vimuth Fernando et al.HPCA 2021 · 11 citations
- XTRA: Unifying Cache Coherence and Concurrency Control for Distributed Transactions in a CXL PodZhijun Yang, Yu Hua, Ming Zhang, Menglei Chen et al.SOSP 2026
- Running Consistent Applications Closer to Users with Radical for Lower LatencyNicolaas Kaashoek, Oleg Aleksandrovich Golev, Austin T. Li, Amit Levy et al.SOSP 2025
- Dorado: Scaling SmartNIC Session Tables on Commodity DDRsHeng Yu, Kai Ren, Jiajun Liang, Baozeng Zhang et al.SIGCOMM 2026
