Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions
Tyler Stennett, Rangeet Pan, Bridget McGinn, Alessandro Orso, Saurabh Sinha
Abstract
Testing is a core activity in software development, and research on its automation has spanned several decades. Most existing approaches focus on generating unit tests for individual methods, validating isolated API endpoints, or targeting user interface (UI) layers. However, for non-API and non-UI tests, automated test generators typically exercise a single focal method. Recent empirical evidence [65] shows a substantial gap between such generated tests and developer-written tests, which often span several focal classes and methods, involve multi-step call sequences, and contain chained assertions-characteristics that current automated approaches fail to capture. To address this gap, we propose generating tests from natural language (NL) descriptions of developer intent. NL provides an expressive and accessible medium for specifying complex test scenarios and functional intent. We present Sakura, the first agent-based framework for generating structurally complex test cases from NL descriptions. Sakura decomposes NL descriptions into structured blocks and processes them using a multi-agent system consisting of a localization agent that grounds test steps in concrete application code via static analysis, a composition agent that synthesizes compilable test code and iteratively refines it using execution feedback, and a supervisor agent that coordinates agent interactions. To evaluate Sakura, we curate a novel dataset of NL test descriptions at three levels of abstraction, reflecting different end-user personas, systematically derived from developer-written tests in Apache Commons projects. Across 20 applications and 1,464 test scenarios, Sakura substantially outperforms off-the-shelf agentic tools such as Gemini CLI instantiated with multiple LLMs. Specifically, Sakura achieves 50-78% higher test compilability and 38-66% higher coverage overlap with ground-truth tests compared to baselines using the same models. Moreover, Sakura paired with small open-source models such as Devstral Small 2 and Qwen3-Coder outperforms Gemini CLI using large proprietary models, while also being more cost-effective.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c88d9005-0ed4-428b-b8e6-11e2bf5edfa0Builds on19
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesMike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li et al.ICLR 2026 · 520 citations
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 163 citations
Related papers
- Generalizing Test Cases for Comprehensive Test Scenario CoverageBinhang Qi, Yun Lin, Xinyi Weng, Chenyan Liu et al.FSE 2026 · 1 citation
- IntentTester: Intent-Driven Multi-agent Framework for Cross-Library Test MigrationYi Gao, Ziyuan Zhang, Xing Hu, Xiaohu Yang et al.FSE 2026
- Industrial Practice of LLM-Based Test Case Carving and Assertion Generation (Experience Paper)Haozhen You, Zhen Dong, Jingjing Wang, Qiang Li et al.ISSTA 2026
- Compiling Large Multi-modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven PerspectiveWeiyu Kong, Yun Lin, Xiwen Teoh, Duc-Minh Nguyen et al.ISSTA 2026 · 1 citation
- Test Intention Guided LLM-Based Unit Test GenerationZifan Nan, Zhaoqiang Guo, Kui Liu, Xin XiaICSE 2025 · 5 citations
