Cell2Doc: ML Pipeline for Generating Documentation in Computational Notebooks
Tamal Mondal, Scott Barnett, Akash Lal, Jyothi Vedurada
摘要
Computational notebooks have become the go-to way for solving data-science problems. While they are designed to combine code and documentation, prior work shows that documentation is largely ignored by the developers because of the manual effort. Automated documentation generation can help, but existing techniques fail to capture algorithmic details and developers often end up editing the generated text to provide more explanation and sub-steps. This paper proposes a novel machine-learning pipeline, Cell2Doc, for code cell documentation in Python data science notebooks. Our approach works by identifying different logical contexts within a code cell, generating documentation for them separately, and finally combining them to arrive at the documentation for the entire code cell. Cell2Doc takes advantage of the capabilities of existing pre-trained language models and improves their efficiency for code cell documentation. We also provide a new benchmark dataset for this task, along with a data-preprocessing pipeline that can be used to create new datasets. We also investigate an appropriate input representation for this task. Our automated evaluation suggests that our best input representation improves the pre-trained model's performance by 2.5x on average. Further, Cell2Doc achieves 1.33x improvement during human evaluation in terms of correctness, informativeness, and readability against the corresponding standalone pretrained model.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Natural Language-Focused Software Engineering via Code-Documentation EquivalenceAryaz Eghbali, Zhongxin Liu, Michael PradelFSE 2026
- NB2P: Generating Data Science Pipelines from Computational NotebooksHaotian Gao, Quang Trung Ta, Tien Tuan Anh Dinh, Nhut-Minh Ho 等ICSE 2026
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 被引用 25 次
- Natural Language to Code Generation in Interactive Data Science NotebooksPengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao 等ACL 2023 · 被引用 17 次
- PyMT5: multi-mode translation of natural language and Python code with transformersColin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy 等EMNLP 2020 · 被引用 24 次
