ParaDox: Eliminating Voltage Margins via Heterogeneous Fault Tolerance
Sam Ainsworth, Lionel Zoubritzky, Alan Mycroft, Timothy M. Jones
Abstract
Providing reliability is becoming a challenge for chip manufacturers, faced with simultaneously trying to improve miniaturization, performance and energy efficiency. This leads to very large margins on voltage and frequency, designed to avoid errors even in the worst case, along with significant hardware expenditure on eliminating voltage spikes and other forms of transient error, causing considerable inefficiency in power consumption and performance.
We flip traditional ideas about reliability and performance around, by exploring the use of error resilience for power and performance gains. ParaMedic is a recent architecture that provides a solution for reliability with low overheads via automatic hardware error recovery. It works by splitting up checking onto many small cores in a heterogeneous multicore system with hardware logging support. However, its design is based on the idea that errors are exceptional. We transform ParaMedic into ParaDox, which shows high performance in both error-intensive and scarce-error scenarios, thus allowing correct execution even when undervolted and overclocked. Evaluation within error-intensive simulation environments confirms the error resilience of ParaDox and the low associated recovery cost. We estimate that compared to a non-resilient system with margins, ParaDox can reduce energy-delay product by 15% through undervolting, while completely recovering from any induced errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e297f0e-1740-4649-ae29-d3d7122660cbCited by top-tier papers3
- Reliability-Aware RunaheadAjeya Naithani, Lieven EeckhoutHPCA 2022 · 2 citations
- MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar ProcessorsZhe Jiang, Minli Liao, Sam Ainsworth, Dean You et al.DAC 2025
- FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time SystemsTinglue Wang, Yiming Li, Wei Tang, Jiapeng Guan et al.DAC 2025
Related papers
- Asymmetric Resilience: Exploiting Task-Level Idempotency for Transient Error Recovery in Accelerator-Based SystemsJingwen Leng, Alper Buyuktosunoglu, Ramon Bertran, Pradip Bose et al.HPCA 2020 · 19 citations
- CARE: Coordinated Augmentation for Elastic Resilience on DRAM Errors in Data CentersJian Chen, Xiaowei Jiang, Ying Zhang, Liyin Liu et al.HPCA 2021 · 9 citations
- RTailor: Parameterizing Soft Error Resilience for Mixed-Criticality Real-Time SystemsShao-Yu Huang, Jianping Zeng, Xuanliang Deng, Sen Wang et al.RTSS 2023 · 11 citations
- Turnpike: Lightweight Soft Error Resilience for In-Order CoresJianping Zeng, Hongjune Kim, Jaejin Lee, Changhee JungMICRO 2021 · 17 citations
- Enabling Software Resilience in GPGPU Applications via Partial Thread ProtectionLishan Yang, Bin Nie, Adwait Jog, Evgenia SmirniICSE 2021 · 25 citations
