LlamaRestTest: Effective REST API Testing with Small Language Models
Myeongsoo Kim, Saurabh Sinha, Alessandro Orso
Abstract
Modern web services rely heavily on REST APIs, typically documented using the OpenAPI specification. The widespread adoption of this standard has resulted in the development of many black-box testing tools that generate tests based on OpenAPI specifications. Although Large Language Models (LLMs) have shown promising test-generation abilities, their application to REST API testing remains mostly unexplored. We present LlamaRestTest, a novel approach that employs two custom LLMs-created by fine-tuning and quantizing the Llama3-8B model using mined datasets of REST API example values and inter-parameter dependencies-to generate realistic test inputs and uncover inter-parameter dependencies during the testing process by analyzing server responses. We evaluated LlamaRestTest on 12 real-world services (including popular services such as Spotify), comparing it against RESTGPT, a GPT-powered specification-enhancement tool, as well as several state-of-the-art REST API testing tools, including RESTler, MoRest, EvoMaster, and ARAT-RL. Our results demonstrate that fine-tuning enables smaller models to outperform much larger models in detecting actionable parameter-dependency rules and generating valid inputs for REST API testing. We also evaluated different tool configurations, ranging from the base Llama3-8B model to fine-tuned versions, and explored multiple quantization techniques, including 2-bit, 4-bit, and 8-bit integer formats. Our study shows that small language models can perform as well as, or better than, large language models in REST API testing, balancing effectiveness and efficiency. Furthermore, LlamaRestTest outperforms state-of-the-art REST API testing tools in code coverage achieved and internal server errors identified, even when those tools use RESTGPT-enhanced specifications. Finally, through an ablation study, we show that each component of LlamaRestTest contributes to its overall performance.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df9ecdbd-2183-42e8-9a83-9505c4bfec02Cited by top-tier papers7
- A Multi-Agent Approach for REST API Testing with Semantic Graphs and LLM-Driven InputsMyeongsoo Kim, Tyler Stennett, Saurabh Sinha, Alessandro OrsoICSE 2025 · 4 citations
- SATORI: Static Test Oracle Generation for REST APIsJuan C. Alonso, Alberto Martin-Lopez, Sergio Segura, Gabriele Bavota et al.ASE 2025 · 2 citations
- Generalizing Test Cases for Comprehensive Test Scenario CoverageBinhang Qi, Yun Lin, Xinyi Weng, Chenyan Liu et al.FSE 2026 · 1 citation
- SAINT: Service-level Integration Test Generation with Program Analysis and LLM-based AgentsRangeet Pan, Raju Pavuluri, Ruikai Huang, Tyler Stennett et al.ICSE 2026 · 1 citation
- RBCTest: Leveraging LLMs to Mine and Verify Oracles of API Response Bodies for RESTful API TestingHieu Huynh, Quoc-Tri Le, Tu Nguyen, Viet Nguyen et al.ICSE 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- The case for 4-bit precision: k-bit Inference Scaling LawsTim Dettmers, Luke ZettlemoyerICML 2023 · 315 citations
Related papers
- Speculate: Generating REST API Specifications using LLMsKrishanu Singh, Kushagra Karar, Abhilash Jindal, Guowei YangFSE 2026
- RESTOR: Automated Test Oracle Generation for RESTful APIs via Reinforcement LearningXun Zhou, Zhen Dong, Mingyu Ren, Qiang Li et al.ISSTA 2026
- Enhancing REST API Testing with NLP TechniquesMyeongsoo Kim, Davide Corradini, Saurabh Sinha, Alessandro Orso et al.ISSTA 2023 · 34 citations
- DeepREST: Automated Test Case Generation for REST APIs Exploiting Deep Reinforcement LearningDavide Corradini, Zeno Montolli, Michele Pasqua, Mariano CeccatoASE 2024 · 13 citations
- MioHint: LLM-Assisted Request Mutation for Whitebox REST API TestingJia Li, Jiacheng Shen, Yuxin Su, Michael R. LyuICSE 2026
