A Closer Look at the Use of Reinforcement Learning for Speeding Up Runtime Verification of Software Tests (Experience Paper)
Shinhae Kim, Saikat Dutta, Owolabi Legunsen
Abstract
Runtime verification (RV) found many bugs by monitoring passing tests against formal specifications (specs), but it is slow. A recent work, Valg, used reinforcement learning (RL) to speed up RV by up to 551.5x or 27 hours. Valg aims to probabilistically monitor most unique traces—sequences of spec-related events like method calls—and monitor fewer redundant ones. But, there is no in-depth study of Valg’s current limits and how to address them. We study Valg on 93 Java open-source projects to answer five unaddressed questions. (i) How much slower is Valg than optimal baselines? Up to 323.8x, or 3.2 hours vs. running tests without RV, and up to 9.6x, or 25.3 minutes vs. a theoretically optimal baseline that monitors only unique traces. (ii) Where is Valg’s time spent? 30.5% on monitoring and 18.4% on signaling events to monitors, on average. (iii) What characterizes code where Valg monitors too many redundant traces or misses unique ones? In 100 cases, 67.8% of redundant traces are due to limitations of Valg’s RL convergence heuristic, and 41.3% of missed unique traces occur when a Valg assumption does not hold. (iv) How much can test non-determinism and RL stochasticity cause monitored unique traces to vary? By 42.5 percentage points (pp) and 12.3pp on average, respectively, but they vary by up to 98pp. (v) How do other off-the-shelf RL algorithms compare with Valg’s? Only two of 11 RL algorithms that we survey are feasible for RV during continuous integration. Both are slower and miss more unique traces than Valg, so custom RL algorithms for RV may be needed. So, despite Valg’s promising results, it has plenty of room to improve. We highlight several exciting future directions on using RL to speed up RV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9338fcf-8600-4ede-aae2-705f9e1ec5f7Builds on5
- An In-Depth Study of Runtime Verification Overheads during Software TestingKevin Guan, Owolabi LegunsenISSTA 2024 · 7 citations
- Faster Explicit-Trace Monitoring-Oriented Programming for Runtime Verification of Software TestsKevin Guan, Marcelo d'Amorim, Owolabi LegunsenOOPSLA 2025 · 7 citations
- Instrumentation-Driven Evolution-Aware Runtime VerificationKevin Guan, Owolabi LegunsenICSE 2025 · 4 citations
- Faster Runtime Verification during Testing via Feedback-Guided Selective MonitoringShinhae Kim, Saikat Dutta, Owolabi LegunsenASE 2025 · 3 citations
- Fine-Grained Analyses for Evolution-Aware Runtime VerificationPengyue Jiang, Kevin Guan, Mahdi Khosravi, Moustafa Ismail et al.ICSE 2026 · 1 citation
Related papers
- Using Reinforcement Learning for Load Testing of Video GamesRosalia Tufano, Simone Scalabrino, Luca Pascarella, Emad Aghajani et al.ICSE 2022 · 37 citations
- Sound and efficient concurrency bug predictionYan Cai, Hao Yun, Jinqiu Wang, Lei Qiao et al.FSE 2021 · 29 citations
- Profiling-Guided Bayesian Optimization of JVM ConfigurationsAbdelrahman Baz, Wing Lam, August ShiISSTA 2026
- Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationAntonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono et al.ICSE 2020 · 81 citations
- Finding Specification Blind Spots via Fuzz TestingRu Ji, Meng XuS&P 2023
