A Closer Look at the Use of Reinforcement Learning for Speeding Up Runtime Verification of Software Tests (Experience Paper)
Shinhae Kim, Saikat Dutta, Owolabi Legunsen
摘要
Runtime verification (RV) found many bugs by monitoring passing tests against formal specifications (specs), but it is slow. A recent work, Valg, used reinforcement learning (RL) to speed up RV by up to 551.5x or 27 hours. Valg aims to probabilistically monitor most unique traces—sequences of spec-related events like method calls—and monitor fewer redundant ones. But, there is no in-depth study of Valg’s current limits and how to address them. We study Valg on 93 Java open-source projects to answer five unaddressed questions. (i) How much slower is Valg than optimal baselines? Up to 323.8x, or 3.2 hours vs. running tests without RV, and up to 9.6x, or 25.3 minutes vs. a theoretically optimal baseline that monitors only unique traces. (ii) Where is Valg’s time spent? 30.5% on monitoring and 18.4% on signaling events to monitors, on average. (iii) What characterizes code where Valg monitors too many redundant traces or misses unique ones? In 100 cases, 67.8% of redundant traces are due to limitations of Valg’s RL convergence heuristic, and 41.3% of missed unique traces occur when a Valg assumption does not hold. (iv) How much can test non-determinism and RL stochasticity cause monitored unique traces to vary? By 42.5 percentage points (pp) and 12.3pp on average, respectively, but they vary by up to 98pp. (v) How do other off-the-shelf RL algorithms compare with Valg’s? Only two of 11 RL algorithms that we survey are feasible for RV during continuous integration. Both are slower and miss more unique traces than Valg, so custom RL algorithms for RV may be needed. So, despite Valg’s promising results, it has plenty of room to improve. We highlight several exciting future directions on using RL to speed up RV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- An In-Depth Study of Runtime Verification Overheads during Software TestingKevin Guan, Owolabi LegunsenISSTA 2024 · 被引用 7 次
- Faster Explicit-Trace Monitoring-Oriented Programming for Runtime Verification of Software TestsKevin Guan, Marcelo d'Amorim, Owolabi LegunsenOOPSLA 2025 · 被引用 7 次
- Instrumentation-Driven Evolution-Aware Runtime VerificationKevin Guan, Owolabi LegunsenICSE 2025 · 被引用 4 次
- Faster Runtime Verification during Testing via Feedback-Guided Selective MonitoringShinhae Kim, Saikat Dutta, Owolabi LegunsenASE 2025 · 被引用 3 次
- Fine-Grained Analyses for Evolution-Aware Runtime VerificationPengyue Jiang, Kevin Guan, Mahdi Khosravi, Moustafa Ismail 等ICSE 2026 · 被引用 1 次
相关 Paper
- Using Reinforcement Learning for Load Testing of Video GamesRosalia Tufano, Simone Scalabrino, Luca Pascarella, Emad Aghajani 等ICSE 2022 · 被引用 37 次
- Sound and efficient concurrency bug predictionYan Cai, Hao Yun, Jinqiu Wang, Lei Qiao 等FSE 2021 · 被引用 29 次
- Profiling-Guided Bayesian Optimization of JVM ConfigurationsAbdelrahman Baz, Wing Lam, August ShiISSTA 2026
- Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationAntonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono 等ICSE 2020 · 被引用 81 次
- Finding Specification Blind Spots via Fuzz TestingRu Ji, Meng XuS&P 2023
