Asymmetric Resilience: Exploiting Task-Level Idempotency for Transient Error Recovery in Accelerator-Based Systems
Jingwen Leng, Alper Buyuktosunoglu, Ramon Bertran, Pradip Bose, Quan Chen, Minyi Guo, Vijay Janapa Reddi
摘要
Accelerators make the task of building systems that are re-silient against transient errors like voltage noise and soft errors hard. Architects integrate accelerators into the system as black box third-party IP components. So a fault in one or more accelerators may threaten the system's reliability if there are no established failure semantics for how an error propagates from the accelerator to the main CPU. Existing solutions that assure system reliability come at the cost of sacrificing accelerator generality, efficiency, and incur significant overhead, even in the absence of errors. To over-come these drawbacks, we examine reliability management of accelerator systems via hardware-software co-design, coupling an efficient architecture design with compiler and run-time support, to cope with transient errors. We introduce asymmetric resilience that architects reliability at the system level, centered around a hardened CPU, rather than at the accelerator level. At runtime, the system exploits task-level idempotency to contain accelerator errors and use memory protection instead of taking checkpoints to mitigate over-heads. We also leverage the fact that errors rarely occur in systems, and exploit the trade-off between error recovery performance and improved error-free performance to enhance system efficiency. Using GPUs, which are at the fore-front of accelerator systems, we demonstrate how our system architecture manages reliability in both integrated and discrete systems, under voltage-noise and soft-error related faults, leading to extremely low overhead (less than 1%) and substantial gains (20% energy savings on average).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong Guo, Chen Zhang, Jingwen Leng, Zihan Liu 等MICRO 2022 · 被引用 109 次
- Ptolemy: Architecture Support for Robust Deep LearningYiming Gan, Yuxian Qiu, Jingwen Leng, Minyi Guo 等MICRO 2020 · 被引用 27 次
- Featherweight Soft Error Resilience for GPUsYida Zhang, Changhee JungMICRO 2022 · 被引用 14 次
- XSched: Preemptive Scheduling for Diverse XPUsWeihang Shen, Mingcong Han, Jialong Liu, Rong Chen 等OSDI 2025 · 被引用 9 次
- Reliability-Aware RunaheadAjeya Naithani, Lieven EeckhoutHPCA 2022 · 被引用 2 次
相关 Paper
- Enabling Software Resilience in GPGPU Applications via Partial Thread ProtectionLishan Yang, Bin Nie, Adwait Jog, Evgenia SmirniICSE 2021 · 被引用 25 次
- G-SEPM: building an accurate and efficient soft error prediction model for GPGPUsHengshan Yue, Xiaohui Wei, Guangli Li, Jianpeng Zhao 等SC 2021 · 被引用 17 次
- Compiler-directed soft error resilience for lightweight GPU register file protectionHongjune Kim, Jianping Zeng, Qingrui Liu, Mohammad Abdel-Majeed 等PLDI 2020 · 被引用 31 次
- FIdelity: Efficient Resilience Analysis Framework for Deep Learning AcceleratorsYi He, Prasanna Balaprakash, Yanjing LiMICRO 2020 · 被引用 82 次
- Turnpike: Lightweight Soft Error Resilience for In-Order CoresJianping Zeng, Hongjune Kim, Jaejin Lee, Changhee JungMICRO 2021 · 被引用 17 次
