Lune

ISSTA2026Top-tier venue

Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions

Tyler Stennett, Rangeet Pan, Bridget McGinn, Alessandro Orso, Saurabh Sinha

2026Year

Abstract

Testing is a core activity in software development, and research on its automation has spanned several decades. Most existing approaches focus on generating unit tests for individual methods, validating isolated API endpoints, or targeting user interface (UI) layers. However, for non-API and non-UI tests, automated test generators typically exercise a single focal method. Recent empirical evidence [65] shows a substantial gap between such generated tests and developer-written tests, which often span several focal classes and methods, involve multi-step call sequences, and contain chained assertions-characteristics that current automated approaches fail to capture. To address this gap, we propose generating tests from natural language (NL) descriptions of developer intent. NL provides an expressive and accessible medium for specifying complex test scenarios and functional intent. We present Sakura, the first agent-based framework for generating structurally complex test cases from NL descriptions. Sakura decomposes NL descriptions into structured blocks and processes them using a multi-agent system consisting of a localization agent that grounds test steps in concrete application code via static analysis, a composition agent that synthesizes compilable test code and iteratively refines it using execution feedback, and a supervisor agent that coordinates agent interactions. To evaluate Sakura, we curate a novel dataset of NL test descriptions at three levels of abstraction, reflecting different end-user personas, systematically derived from developer-written tests in Apache Commons projects. Across 20 applications and 1,464 test scenarios, Sakura substantially outperforms off-the-shelf agentic tools such as Gemini CLI instantiated with multiple LLMs. Specifically, Sakura achieves 50-78% higher test compilability and 38-66% higher coverage overlap with ground-truth tests compared to baselines using the same models. Moreover, Sakura paired with small open-source models such as Devstral Small 2 and Qwen3-Coder outperforms Gemini CLI using large proprietary models, while also being more cost-effective.

CCS Concepts: • Software and its engineering → Software testing and debugging.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c88d9005-0ed4-428b-b8e6-11e2bf5edfa0

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines