Doc2OracLL: Investigating the Impact of Documentation on LLM-Based Test Oracle Generation
Soneya Binta Hossain, Raygan Taylor, Matthew B. Dwyer
Abstract
Code documentation is a critical artifact of software development, bridging human understanding and machinereadable code. Beyond aiding developers in code comprehension and maintenance, documentation also plays a critical role in automating various software engineering tasks, such as test oracle generation (TOG). In Java, Javadoc comments offer structured, natural language documentation embedded directly within the source code, typically describing functionality, usage, parameters, return values, and exceptional behavior. While prior research has explored the use of Javadoc comments in TOG alongside other information, such as the method under test, their potential as a stand-alone input source, the most relevant Javadoc components, and guidelines for writing effective Javadoc comments for automating TOG remain less explored.. In this study, we investigate the impact of Javadoc comments on TOG through a comprehensive analysis. We begin by fine-tuning 10 large language models using three different prompt pairs to assess the role of Javadoc comments alongside other contextual information. Next, we systematically analyze the impact of different Javadoc comment's components on TOG. To evaluate the generalizability of Javadoc comments from various sources, we also generate them using the GPT-3.5 model. We perform a thorough bug detection study using Defects4J dataset to understand their role in real-world bug detection. Our results show that incorporating Javadoc comments improves the accuracy of test oracles in most cases, aligning closely with ground truth. We find that Javadoc comments alone can achieve comparable or even better performance when using the implementation of MUT. Additionally, we identify that the description and the return tag are the most valuable components for TOG. Finally, our approach, when using only Javadoc comments, detects between 19% and 94% more real-world bugs in Defects4J than prior methods, establishing a new state-of-the-art. To further guide developers in writing effective documentation, we conduct a detailed qualitative study on when Javadoc comments are helpful or harmful for TOG. CCS Concepts: • Software and its engineering → Software testing and debugging; Documentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9c05813-13ab-4285-906b-2d6ef7aac8ceCited by top-tier papers6
- Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue RepairKai Huang, Jian Zhang, Xiaofei Xie, Chunyang ChenASE 2025 · 5 citations
- Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit TestsJunda Zhao, Shurui Zhou, Eldan CohenISSTA 2026 · 1 citation
- Inside Out: Uncovering How Comment Internalization Steers LLMs for Better or WorseAaron Imani, Mohammad Moshirpour, Iftekhar AhmedICSE 2026
- Code-MUE: Measuring Code LLMs’ Uncertainty through Execution-Based Semantic Interaction GraphsXiaoning Ren, Yinxing Xue, Lei Ma, Yuheng HuangISSTA 2026
- LogicHunter: Testing LLM Agent Frameworks with an Agentic OracleMinghui Long, Yanjie Zhao, Haoyu WangISSTA 2026
Builds on11
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 1,026 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- Software documentation: the practitioners' perspectiveEmad Aghajani, Csaba Nagy, Mario Linares-Vásquez, Laura Moreno et al.ICSE 2020 · 112 citations
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 103 citations
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota et al.ICSE 2020 · 96 citations
Related papers
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 12 citations
- Neural-Based Test Oracle Generation: A Large-Scale Evaluation and Lessons LearnedSoneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum et al.FSE 2023 · 30 citations
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 92 citations
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst et al.ASE 2025 · 3 citations
