Industrial Practice of LLM-Based Test Case Carving and Assertion Generation (Experience Paper)
Haozhen You, Zhen Dong, Jingjing Wang, Qiang Li, Xin Peng
摘要
Enterprise regression testing for microservice systems is often constrained by incomplete or outdated documentation. In practice, QA engineers frequently rely on real execution traffic to reconstruct business scenarios; however, turning raw traffic into replayable regression tests with stable validation logic remains labor-intensive and error-prone.
This paper presents NL2Test, an end-to-end approach and tool that generates executable API regression tests from (i) a natural-language scenario description and (ii) a traffic capture recorded while executing the scenario. NL2Test addresses two coupled tasks: test case carving, which extracts a minimal replayable request sequence and reconstructs data dependencies so that dynamic values are bound from their responses rather than hard-coded; and assertion generation, which produces assertions aligned with business intent while avoiding non-deterministic fields and hallucinated paths. To improve reliability, NL2Test uses LLMs for semantic interpretation and constrained code synthesis, and uses deterministic algorithms for request filtering, dependency confirmation via value consistency, and assertion-path validation.
We evaluate NL2Test on 51 industrial regression scenarios extracted from a large consumer-facing Internet company. NL2Test achieves an exact-match rate of 82.4% (42/51), and produces a functionally usable draft in 98.0% (50/51) of scenarios when allowing minor post-edits. In a 9-month production deployment starting in March 2025, NL2Test generated 3,196 test cases with an overall code adoption rate of 85.4%. These results indicate that traffic-grounded generation with deterministic guardrails can substantially reduce manual effort while improving regression automation in complex microservice environments.
CCS Concepts: • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Time-travel testing of Android appsZhen Dong, Marcel Böhme, Lucia Cojocaru, Abhik RoychoudhuryICSE 2020 · 被引用 104 次
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
- Understanding and finding system setting-related defects in Android appsJingling Sun, Ting Su, Junxin Li, Zhen Dong 等ISSTA 2021 · 被引用 35 次
- Carving UI Tests to Generate API Tests and API SpecificationRahulkrishna Yandrapally, Saurabh Sinha, Rachel Tzoref-Brill, Ali MesbahICSE 2023 · 被引用 21 次
- Detecting and fixing data loss issues in Android appsWunan Guo, Zhen Dong, Liwei Shen, Wei Tian 等ISSTA 2022 · 被引用 17 次
相关 Paper
- Sakura: An Approach for Generating Complex Tests from Natural Language Test DescriptionsTyler Stennett, Rangeet Pan, Bridget McGinn, Alessandro Orso 等ISSTA 2026
- LlamaRestTest: Effective REST API Testing with Small Language ModelsMyeongsoo Kim, Saurabh Sinha, Alessandro OrsoFSE 2025 · 被引用 9 次
- Enhancing REST API Testing with NLP TechniquesMyeongsoo Kim, Davide Corradini, Saurabh Sinha, Alessandro Orso 等ISSTA 2023 · 被引用 34 次
- Uncovering Business Logic Bugs via Semantics-Driven Unit Test Generation (Experience Paper)Chen Yang, Junjie ChenISSTA 2026
- E-Test: E'er-Improving Test SuitesKetai Qiu, Luca Di Grazia, Leonardo Mariani, Mauro PezzèICSE 2026
