Assessing and Restoring Reproducibility of Jupyter Notebooks
Jiawei Wang, Tzu-yang Kuo, Li Li, Andreas Zeller
Abstract
Jupyter notebooks-documents that contain live code, equations, visualizations, and narrative text-now are among the most popular means to compute, present, discuss and disseminate scientific findings. In principle, Jupyter notebooks should easily allow to reproduce and extend scientific computations and their findings; but in practice, this is not the case. The individual code cells in Jupyter notebooks can be executed in any order, with identifier usages preceding their definitions and results preceding their computations. In a sample of 936 published notebooks that would be executable in principle, we found that 73% of them would not be reproducible with straightforward approaches, requiring humans to infer (and often guess) the order in which the authors created the cells. In this paper, we present an approach to (1) automatically satisfy dependencies between code cells to reconstruct possible execution orders of the cells; and (2) instrument code cells to mitigate the impact of non-reproducible statements (i.e., random functions) in Jupyter notebooks. Our Osiris prototype takes a notebook as input and outputs the possible execution schemes that reproduce the exact notebook results. In our sample, Osiris was able to reconstruct such schemes for 82.23% of all executable notebooks, which has more than three times better than the state-of-the-art; the resulting reordered code is valid program code and thus available for further testing and analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10bba54c-666c-4cbb-80d9-4c93b7138bceCited by top-tier papers14
- Exploring how deprecated Python library APIs are (not) handledJiawei Wang, Li Li, Kui Liu, Haipeng CaiFSE 2020 · 52 citations
- Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task DecompositionMajeed Kazemitabaar, Jack Williams, Ian Drosos, Tovi Grossman et al.UIST 2024 · 49 citations
- Restoring Execution Environments of Jupyter NotebooksJiawei Wang, Li Li, Andreas ZellerICSE 2021 · 49 citations
- How Do Analysts Understand and Verify AI-Assisted Data Analyses?Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang et al.CHI 2024 · 36 citations
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 25 citations
Related papers
- Fine-Grained Lineage for Safer Notebook InteractionsStephen Macke, Aditya G. Parameswaran, Hongpu Gong, Doris Jung Lin Lee et al.VLDB 2021 · 46 citations
- Bolt-on, Compact, and Rapid Program Slicing for Notebooks [Scalable Data Science]Shreya Shankar, Stephen Macke, Sarah E. Chasins, Andrew Head et al.VLDB 2022 · 17 citations
- NB2P: Generating Data Science Pipelines from Computational NotebooksHaotian Gao, Quang Trung Ta, Tien Tuan Anh Dinh, Nhut-Minh Ho et al.ICSE 2026
- ElasticNotebook: Enabling Live Migration for Computational NotebooksZhaoheng Li, Pranav Gor, Rahul Prabhu, Hui Yu et al.VLDB 2024 · 12 citations
- Automated Modernization of Machine Learning Engineering Notebooks for ReproducibilityBihui Jin, Kaiyuan Wang, Pengyu NieISSTA 2026
