Dvé: Improving DRAM Reliability and Performance On-Demand via Coherent Replication
Adarsh Patil, Vijay Nagarajan, Rajeev Balasubramonian, Nicolai Oswald
Abstract
As technologies continue to shrink, memory system failure rates have increased, demanding support for stronger forms of reliability. In this work, we take inspiration from the two-tier approach that decouples correction from detection and explore a novel extrapolation. We propose Dvé, a hardware-driven replication mechanism where data blocks are replicated in 2 different sockets across a cache-coherent NUMA system. Each data block is also accompanied by a code with strong error detection capabilities so that when an error is detected, correction is performed using the replica. Such an organization has the advantage of offering two independent points of access to data which enables: (a) strong error correction that can recover from a range of faults affecting any of the components in the memory, upto and including the memory controller, and (b) higher performance by providing another nearer point of memory access. Dvé realizes both of these benefits via Coherent Replication, a technique that builds on top of existing cache coherence protocols for not only keeping the replicas in sync for reliability, but also to provide coherent access to the replicas during fault-free operation for performance. Dvé can flexibly provide these benefits on-demand by simply using the provisioned memory capacity which, as reported in recent studies, is often underutilized in today’s systems. Thus, Dvé introduces a unique design point that offers higher reliability and performance for workloads that do not require the entire memory capacity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0846ec1c-1b0b-4608-a592-ea3e2b7454f6Cited by top-tier papers3
- DyLeCT: Achieving Huge-page-like Translation Performance for Hardware-compressed MemoryGagandeep Panwar, Muhammad Laghari, Esha Choukse, Xun JianISCA 2024 · 4 citations
- Polymorphic Error CorrectionEvgeny Manzhosov, Simha SethumadhavanMICRO 2024 · 1 citation
- Dorado: Clustered Hardware Cache Coherence for 1,000+ CoresJovan Stojkovic, Abraham Farrell, Gerasimos Gerogiannis, Zhangxiaowen Gong et al.ISCA 2026 · 1 citation
Builds on1
Related papers
- UniMem: Redesigning Disaggregated Memory within A Unified Local-Remote Memory HierarchyYijie Zhong, Minqiang Zhou, Zhirong Shen, Jiwu ShuUSENIX ATC 2024 · 7 citations
- CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersJian Chen, Xiaowei Jiang, Ying Zhang, Liyin Liu et al.HPCA 2021 · 9 citations
- XTRA: Unifying Cache Coherence and Concurrency Control for Distributed Transactions in a CXL PodZhijun Yang, Yu Hua, Ming Zhang, Menglei Chen et al.SOSP 2026
- PaCaR: Improved Buffered I/O Locality on NUMA Systems with Page Cache ReplicationJérôme Coquisart, Julien Sopena, Redha GouicemEuroSys 2026
- Unity ECC: Unified Memory Protection Against Bit and Chip ErrorsDongwhee Kim, Jaeyoon Lee, Wonyeong Jung, Michael B. Sullivan et al.SC 2023 · 23 citations
