Cell2Doc: ML Pipeline for Generating Documentation in Computational Notebooks
Tamal Mondal, Scott Barnett, Akash Lal, Jyothi Vedurada
Abstract
Computational notebooks have become the go-to way for solving data-science problems. While they are designed to combine code and documentation, prior work shows that documentation is largely ignored by the developers because of the manual effort. Automated documentation generation can help, but existing techniques fail to capture algorithmic details and developers often end up editing the generated text to provide more explanation and sub-steps. This paper proposes a novel machine-learning pipeline, Cell2Doc, for code cell documentation in Python data science notebooks. Our approach works by identifying different logical contexts within a code cell, generating documentation for them separately, and finally combining them to arrive at the documentation for the entire code cell. Cell2Doc takes advantage of the capabilities of existing pre-trained language models and improves their efficiency for code cell documentation. We also provide a new benchmark dataset for this task, along with a data-preprocessing pipeline that can be used to create new datasets. We also investigate an appropriate input representation for this task. Our automated evaluation suggests that our best input representation improves the pre-trained model's performance by 2.5x on average. Further, Cell2Doc achieves 1.33x improvement during human evaluation in terms of correctness, informativeness, and readability against the corresponding standalone pretrained model.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1ef63c1e-a7b2-46db-83e5-bc4323b30931Related papers
- Natural Language-Focused Software Engineering via Code-Documentation EquivalenceAryaz Eghbali, Zhongxin Liu, Michael PradelFSE 2026
- NB2P: Generating Data Science Pipelines from Computational NotebooksHaotian Gao, Quang Trung Ta, Tien Tuan Anh Dinh, Nhut-Minh Ho et al.ICSE 2026
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 25 citations
- Natural Language to Code Generation in Interactive Data Science NotebooksPengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao et al.ACL 2023 · 17 citations
- PyMT5: multi-mode translation of natural language and Python code with transformersColin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy et al.EMNLP 2020 · 24 citations
