Dr. DNA: Combating Silent Data Corruptions in Deep Learning using Distribution of Neuron Activations
Dongning Ma, Fan Fred Lin, Alban Desmaison, Joel Coburn, Daniel Moore, Sriram Sankar, Xun Jiao
摘要
Deep neural networks (DNNs) have been widely-adopted in various safety-critical applications such as computer vision and autonomous driving. However, as technology scales and applications diversify, coupled with the increasing heterogeneity of underlying hardware architectures, silent data corruption (SDC) has been emerging as a pronouncing threat to the reliability of DNNs. Recent reports from industry hyperscalars underscore the difficulty in addressing SDC due to their "stealthy" nature and elusive manifestation. In this paper, we propose Dr. DNA, a novel approach to enhance the reliability of DNN systems by detecting and mitigating SDCs. Specifically, we formulate and extract a set of unique SDC signatures from the Distribution of Neuron Activations (DNA), based on which we propose early-stage detection and mitigation of SDCs during DNN inference. We perform an extensive evaluation across 3 vision tasks, 5 different datasets, and 10 different models, under 4 different error models. Results show that Dr. DNA achieves 100% SDC detection rate for most cases, 95% detection rate on average and >90% detection rate across all cases, representing 20% - 70% improvement over baselines. Dr. DNA can also mitigate the impact of SDCs by effectively recovering DNN model performance with <1% memory overhead and <2.5% latency overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Understanding Silent Data Corruption in LLM TrainingJeffrey Jian Ma, Hengzhi Pei, Leonard Lausen, George KarypisACL 2025 · 被引用 20 次
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park 等ICLR 2026
- Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU HoursKinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen 等OSDI 2026
相关 Paper
- DeepDyve: Dynamic Verification for Deep Neural NetworksYu Li, Min Li, Bo Luo, Ye Tian 等CCS 2020 · 被引用 28 次
- SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale DatacentersYixuan Mei, Shreya Varshini, Harish Dattatraya Dixit, Sriram Sankar 等ASPLOS 2026
- Detection of Out-of-Distribution Samples Using Binary Neuron Activation PatternsBartlomiej Olber, Krystian Radlak, Adam Popowicz, Michal Szczepankiewicz 等CVPR 2023
- Misbehaviour prediction for autonomous driving systemsAndrea Stocco, Michael Weiss, Marco Calzana, Paolo TonellaICSE 2020 · 被引用 138 次
- AegisDNN: Dependable and Timely Execution of DNN Tasks with SGXYecheng Xiang, Yidi Wang, Hyunjong Choi, Mohsen Karimi 等RTSS 2021 · 被引用 23 次
