Lune

ISSTA2026顶会

Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions

Tyler Stennett, Rangeet Pan, Bridget McGinn, Alessandro Orso, Saurabh Sinha

2026年份

摘要

Testing is a core activity in software development, and research on its automation has spanned several decades. Most existing approaches focus on generating unit tests for individual methods, validating isolated API endpoints, or targeting user interface (UI) layers. However, for non-API and non-UI tests, automated test generators typically exercise a single focal method. Recent empirical evidence [65] shows a substantial gap between such generated tests and developer-written tests, which often span several focal classes and methods, involve multi-step call sequences, and contain chained assertions-characteristics that current automated approaches fail to capture. To address this gap, we propose generating tests from natural language (NL) descriptions of developer intent. NL provides an expressive and accessible medium for specifying complex test scenarios and functional intent. We present Sakura, the first agent-based framework for generating structurally complex test cases from NL descriptions. Sakura decomposes NL descriptions into structured blocks and processes them using a multi-agent system consisting of a localization agent that grounds test steps in concrete application code via static analysis, a composition agent that synthesizes compilable test code and iteratively refines it using execution feedback, and a supervisor agent that coordinates agent interactions. To evaluate Sakura, we curate a novel dataset of NL test descriptions at three levels of abstraction, reflecting different end-user personas, systematically derived from developer-written tests in Apache Commons projects. Across 20 applications and 1,464 test scenarios, Sakura substantially outperforms off-the-shelf agentic tools such as Gemini CLI instantiated with multiple LLMs. Specifically, Sakura achieves 50-78% higher test compilability and 38-66% higher coverage overlap with ground-truth tests compared to baselines using the same models. Moreover, Sakura paired with small open-source models such as Devstral Small 2 and Qwen3-Coder outperforms Gemini CLI using large proprietary models, while also being more cost-effective.

CCS Concepts: • Software and its engineering → Software testing and debugging.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖