Dr. DNA: Combating Silent Data Corruptions in Deep Learning using Distribution of Neuron Activations
Dongning Ma, Fan Fred Lin, Alban Desmaison, Joel Coburn, Daniel Moore, Sriram Sankar, Xun Jiao
Abstract
Deep neural networks (DNNs) have been widely-adopted in various safety-critical applications such as computer vision and autonomous driving. However, as technology scales and applications diversify, coupled with the increasing heterogeneity of underlying hardware architectures, silent data corruption (SDC) has been emerging as a pronouncing threat to the reliability of DNNs. Recent reports from industry hyperscalars underscore the difficulty in addressing SDC due to their "stealthy" nature and elusive manifestation. In this paper, we propose Dr. DNA, a novel approach to enhance the reliability of DNN systems by detecting and mitigating SDCs. Specifically, we formulate and extract a set of unique SDC signatures from the Distribution of Neuron Activations (DNA), based on which we propose early-stage detection and mitigation of SDCs during DNN inference. We perform an extensive evaluation across 3 vision tasks, 5 different datasets, and 10 different models, under 4 different error models. Results show that Dr. DNA achieves 100% SDC detection rate for most cases, 95% detection rate on average and >90% detection rate across all cases, representing 20% - 70% improvement over baselines. Dr. DNA can also mitigate the impact of SDCs by effectively recovering DNN model performance with <1% memory overhead and <2.5% latency overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 09e952a3-cd9d-4e92-a705-9b961af60a17Cited by top-tier papers3
- Understanding Silent Data Corruption in LLM TrainingJeffrey Jian Ma, Hengzhi Pei, Leonard Lausen, George KarypisACL 2025 · 20 citations
- SpareTrain: Fault-Tolerant LLM Training via Low-Cost Dual Modular RedundancyRihae Park, Yeonjae Kim, Seung Yul Lee, Yeonhong Park et al.ICLR 2026
- Safeguarding LLM Training at Scale: Online SDC Detection and Insights from 35 Million GPU HoursKinman Lei, Liyan Zheng, Xiang Li, Hongmin Chen et al.OSDI 2026
Related papers
- DeepDyve: Dynamic Verification for Deep Neural NetworksYu Li, Min Li, Bo Luo, Ye Tian et al.CCS 2020 · 28 citations
- SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale DatacentersYixuan Mei, Shreya Varshini, Harish Dattatraya Dixit, Sriram Sankar et al.ASPLOS 2026
- Detection of Out-of-Distribution Samples Using Binary Neuron Activation PatternsBartlomiej Olber, Krystian Radlak, Adam Popowicz, Michal Szczepankiewicz et al.CVPR 2023
- Misbehaviour prediction for autonomous driving systemsAndrea Stocco, Michael Weiss, Marco Calzana, Paolo TonellaICSE 2020 · 138 citations
- AegisDNN: Dependable and Timely Execution of DNN Tasks with SGXYecheng Xiang, Yidi Wang, Hyunjong Choi, Mohsen Karimi et al.RTSS 2021 · 23 citations
