Industrial Practice of LLM-Based Test Case Carving and Assertion Generation (Experience Paper)
Haozhen You, Zhen Dong, Jingjing Wang, Qiang Li, Xin Peng
Abstract
Enterprise regression testing for microservice systems is often constrained by incomplete or outdated documentation. In practice, QA engineers frequently rely on real execution traffic to reconstruct business scenarios; however, turning raw traffic into replayable regression tests with stable validation logic remains labor-intensive and error-prone.
This paper presents NL2Test, an end-to-end approach and tool that generates executable API regression tests from (i) a natural-language scenario description and (ii) a traffic capture recorded while executing the scenario. NL2Test addresses two coupled tasks: test case carving, which extracts a minimal replayable request sequence and reconstructs data dependencies so that dynamic values are bound from their responses rather than hard-coded; and assertion generation, which produces assertions aligned with business intent while avoiding non-deterministic fields and hallucinated paths. To improve reliability, NL2Test uses LLMs for semantic interpretation and constrained code synthesis, and uses deterministic algorithms for request filtering, dependency confirmation via value consistency, and assertion-path validation.
We evaluate NL2Test on 51 industrial regression scenarios extracted from a large consumer-facing Internet company. NL2Test achieves an exact-match rate of 82.4% (42/51), and produces a functionally usable draft in 98.0% (50/51) of scenarios when allowing minor post-edits. In a 9-month production deployment starting in March 2025, NL2Test generated 3,196 test cases with an overall code adoption rate of 85.4%. These results indicate that traffic-grounded generation with deterministic guardrails can substantially reduce manual effort while improving regression automation in complex microservice environments.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d28e5f23-3bc7-43e3-bc16-35595634e244Builds on10
- Time-travel testing of Android appsZhen Dong, Marcel Böhme, Lucia Cojocaru, Abhik RoychoudhuryICSE 2020 · 104 citations
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 92 citations
- Understanding and finding system setting-related defects in Android appsJingling Sun, Ting Su, Junxin Li, Zhen Dong et al.ISSTA 2021 · 35 citations
- Carving UI Tests to Generate API Tests and API SpecificationRahulkrishna Yandrapally, Saurabh Sinha, Rachel Tzoref-Brill, Ali MesbahICSE 2023 · 21 citations
- Detecting and fixing data loss issues in Android appsWunan Guo, Zhen Dong, Liwei Shen, Wei Tian et al.ISSTA 2022 · 17 citations
Related papers
- Sakura: An Approach for Generating Complex Tests from Natural Language Test DescriptionsTyler Stennett, Rangeet Pan, Bridget McGinn, Alessandro Orso et al.ISSTA 2026
- LlamaRestTest: Effective REST API Testing with Small Language ModelsMyeongsoo Kim, Saurabh Sinha, Alessandro OrsoFSE 2025 · 9 citations
- Enhancing REST API Testing with NLP TechniquesMyeongsoo Kim, Davide Corradini, Saurabh Sinha, Alessandro Orso et al.ISSTA 2023 · 34 citations
- Uncovering Business Logic Bugs via Semantics-Driven Unit Test Generation (Experience Paper)Chen Yang, Junjie ChenISSTA 2026
- E-Test: E'er-Improving Test SuitesKetai Qiu, Luca Di Grazia, Leonardo Mariani, Mauro PezzèICSE 2026
