Cockpit: A Practical Debugging Tool for the Training of Deep Neural Networks
Frank Schneider, Felix Dangel, Philipp Hennig
摘要
When engineers train deep learning models, they are very much 'flying blind'. Commonly used methods for real-time training diagnostics, such as monitoring the train/test loss, are limited. Assessing a network's training process solely through these performance indicators is akin to debugging software without access to internal states through a debugger. To address this, we present Cockpit, a collection of instruments that enable a closer look into the inner workings of a learning machine, and a more informative and meaningful status report for practitioners. It facilitates the identification of learning phases and failure modes, like ill-chosen hyperparameters. These instruments leverage novel higher-order information about the gradient distribution and curvature, which has only recently become efficiently accessible. We believe that such a debugging tool, which we open-source for PyTorch, is a valuable help in troubleshooting the training process. By revealing new insights, it also more generally contributes to explainability and interpretability of deep nets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PCcheck: Persistent Concurrent Checkpointing for MLFoteini Strati, Michal Friedman, Ana KlimovicASPLOS 2025 · 被引用 11 次
- Per-example Gradients: a New Frontier for Understanding and Improving OptimizersVincent Roulet, Atish AgarwalaICML 2026 · 被引用 2 次
- BTTackler: A Diagnosis-based Framework for Efficient Deep Learning Hyperparameter OptimizationZhongyi Pei, Zhiyao Cen, Yipeng Huang, Chen Wang 等KDD 2024 · 被引用 1 次
- DeepDebugger: An Interactive Time-Travelling Debugging Approach for Deep ClassifiersXianglin Yang, Yun Lin, Yifan Zhang, Linpeng Huang 等FSE 2023 · 被引用 1 次
它引用的顶会 Paper7
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 被引用 199 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang 等ICML 2020 · 被引用 122 次
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg 等ICML 2021 · 被引用 78 次
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based OptimizationSatrajit ChatterjeeICLR 2020 · 被引用 60 次
相关 Paper
- Hindsight Logging for Model TrainingRolando Garcia, Eric Liu, Vikram Sreekanti, Bobby Yan 等VLDB 2021 · 被引用 10 次
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 被引用 62 次
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 被引用 114 次
- Skyline: Interactive In-Editor Computational Performance Profiling for Deep Neural Network TrainingGeoffrey X. Yu, Tovi Grossman, Gennady PekhimenkoUIST 2020 · 被引用 16 次
- DeepDiagnosis: Automatically Diagnosing Faults and Recommending Actionable Fixes in Deep Learning ProgramsMohammad Wardat, Breno Dantas Cruz, Wei Le, Hridesh RajanICSE 2022 · 被引用 46 次
