Hindsight Logging for Model Training
Rolando Garcia, Eric Liu, Vikram Sreekanti, Bobby Yan, Anusha Dandamudi, Joseph Gonzalez, Joseph M. Hellerstein, Koushik Sen
摘要
In modern Machine Learning, model training is an iterative, experimental process that can consume enormous computation resources and developer time. To aid in that process, experienced model developers log and visualize program variables during training runs. Exhaustive logging of all variables is infeasible, so developers are left to choose between slowing down training via extensive conservative logging, or letting training run fast via minimalist optimistic logging that may omit key information. As a compromise, optimistic logging can be accompanied by program checkpoints; this allows developers to add log statements post-hoc, and "replay" desired log statements from checkpoint-a process we refer to as hindsight logging. Unfortunately, hindsight logging raises tricky problems in data management and software engineering. Done poorly, hindsight logging can waste resources and generate technical debt embodied in multiple variants of training code. In this paper, we present methodologies for efficient and effective logging practices for model training, with a focus on techniques for hindsight logging. Our goal is for experienced model developers to learn and adopt these practices. To make this easier, we provide an open-source suite of tools for Fast Low-Overhead Recovery (flor) that embodies our design across three tasks: (i) efficient background logging in Python, (ii) adaptable periodic checkpointing, and (iii) an instrumentation library that codifies hindsight logging for efficient and automatic record-replay of model-training. Model developers can use each flor tool separately as they see fit, or they can use flor in hands-free mode, entrusting it to instrument their code end-to-end for efficient record-replay. Our solutions leverage techniques from physiological transaction logs and recovery in database systems. Evaluations on modern ML benchmarks demonstrate that flor can produce fast checkpointing with small user-specifiable overheads (e.g. 7%), and still provide hindsight log replay times orders of magnitude faster than restarting training from scratch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- "We Have No Idea How Models will Behave in Production until Production": How Engineers Operationalize Machine LearningShreya Shankar, Rolando Garcia, Joseph M. Hellerstein, Aditya G. ParameswaranCSCW 2024 · 被引用 27 次
- Towards Observability for Production Machine Learning Pipelines [Vision]Shreya Shankar, Aditya G. ParameswaranVLDB 2022 · 被引用 21 次
- ElasticNotebook: Enabling Live Migration for Computational NotebooksZhaoheng Li, Pranav Gor, Rahul Prabhu, Hui Yu 等VLDB 2024 · 被引用 12 次
- CHEX: Multiversion Replay with Ordered CheckpointsNaga Nithin Manne, Shilvi Satpati, Tanu Malik, Amitabha Bagchi 等VLDB 2022 · 被引用 9 次
它引用的顶会 Paper1
相关 Paper
- Cockpit: A Practical Debugging Tool for the Training of Deep Neural NetworksFrank Schneider, Felix Dangel, Philipp HennigNeurIPS 2021 · 被引用 14 次
- Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model TrainingZichen Wang, Hongliang Li, Jie Wu, Zhewen Xu 等INFOCOM 2026
- PCcheck: Persistent Concurrent Checkpointing for MLFoteini Strati, Michal Friedman, Ana KlimovicASPLOS 2025 · 被引用 11 次
- Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training FailuresTanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev 等EuroSys 2024 · 被引用 23 次
- Relight: Simple User-Level Checkpointing and Fast-Forward Replay for Distributed Task-Based SystemsElliott Slaughter, Rupanshu Soi, Michael Bauer, Alex AikenOOPSLA 2026 · 被引用 1 次
