CHEX: Multiversion Replay with Ordered Checkpoints
Naga Nithin Manne, Shilvi Satpati, Tanu Malik, Amitabha Bagchi, Ashish Gehani, Amitabh Chaudhary
Abstract
In scientific computing and data science disciplines, it is often necessary to share application workflows and repeat results. Current tools containerize application workflows, and share the resulting container for repeating results. These tools, due to containerization, do improve sharing of results. However, they do not improve the efficiency of replay. In this paper, we present the multiversion replay problem, which arises when multiple versions of an application are containerized, and each version must be replayed to repeat results. To avoid executing each version separately, we develop CHEX , which checkpoints program state and determines when it is permissible to reuse program state across versions. It does so using system call-based execution lineage. Our capability to identify common computations across versions enables us to consider optimizing replay using an in-memory cache, based on a checkpoint-restore-switch system. We show the multiversion replay problem is NP-hard, and propose efficient heuristics for it. CHEX reduces overall replay time by sharing common computations but avoids storing a large number of checkpoints. We demonstrate that CHEX maintains lightweight package sharing, and improves the total time of multiversion replay by 50% on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51ecc5b0-b216-458a-8031-55abe9045026Cited by top-tier papers4
- ElasticNotebook: Enabling Live Migration for Computational NotebooksZhaoheng Li, Pranav Gor, Rahul Prabhu, Hui Yu et al.VLDB 2024 · 12 citations
- Kishu: Time-Traveling for Computational NotebooksZhaoheng Li, Supawit Chockchowwat, Areet Sheth, Yongjoo Park et al.VLDB 2025 · 11 citations
- Enhancing Computational Notebooks with Code+Data Space VersioningHanxi Fang, Supawit Chockchowwat, Hari Sundaram, Yongjoo ParkCHI 2025 · 6 citations
- Kondo: Efficient Provenance-Driven Data DebloatingAniket Modi, Rohan Tikmany, Tanu Malik, Raghavan Komondoor et al.ICDE 2024 · 4 citations
Builds on2
Related papers
- AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency AnalysisXiang Fu, Weiping Zhang, Shiman Meng, Xin Huang et al.SC 2024 · 1 citation
- Relight: Simple User-Level Checkpointing and Fast-Forward Replay for Distributed Task-Based SystemsElliott Slaughter, Rupanshu Soi, Michael Bauer, Alex AikenOOPSLA 2026 · 1 citation
- Reproducible ContainersOmar S. Navarro Leija, Kelly Shiptoski, Ryan G. Scott, Baojun Wang et al.ASPLOS 2020 · 21 citations
- LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning SystemsArnab Phani, Benjamin Rath, Matthias BoehmSIGMOD 2021 · 30 citations
- Optimizing Machine Learning Workloads in Collaborative EnvironmentsBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl et al.SIGMOD 2020 · 22 citations
