Policy Testing with MDPFuzz (Replicability Study)
Quentin Mazouni, Helge Spieker, Arnaud Gotlieb, Mathieu Acher
摘要
In recent years, following tremendous achievements in Reinforcement Learning, a great deal of interest has been devoted to ML models for sequential decision-making. Together with these scientific breakthroughs/advances, research has been conducted to develop automated functional testing methods for finding faults in black-box Markov decision processes. Pang et al. (ISSTA 2022) presented a black-box fuzz testing framework called MDPFuzz. The method consists of a fuzzer whose main feature is to use Gaussian Mixture Models (GMMs) to compute coverage of the test inputs as the likelihood to have already observed their results. This guidance through coverage evaluation aims at favoring novelty during testing and fault discovery in the decision model. Pang et al. evaluated their work with four use cases, by comparing the number of failures found after twelve-hour testing campaigns with or without the guidance of the GMMs (ablation study). In this paper, we verify some of the key findings of the original paper and explore the limits of MDPFuzz through reproduction and replication. We re-implemented the proposed methodology and evaluated our replication in a large-scale study that extends the original four use cases with three new ones. Furthermore, we compare MDPFuzz and its ablated counterpart with a random testing baseline. We also assess the effectiveness of coverage guidance for different parameters, something that has not been done in the original evaluation. Despite this parameter analysis and unlike Pang et al. 's original conclusions, we find that in most cases, the aforementioned ablated Fuzzer outperforms MDPFuzz, and conclude that the coverage model proposed does not lead to finding more faults. CCS Concepts • Software and its engineering → Software testing and debugging; • Computing methodologies → Reinforcement learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Learning Generalized Relational Heuristic Networks for Model-Agnostic PlanningRushang Karia, Siddharth SrivastavaAAAI 2021 · 被引用 49 次
- MDPFuzz: testing models solving Markov decision processesQi Pang, Yuanyuan Yuan, Shuai WangISSTA 2022 · 被引用 37 次
- Many-Objective Reinforcement Learning for Online Testing of DNN-Enabled SystemsFitash Ul Haq, Donghwan Shin, Lionel C. BriandICSE 2023 · 被引用 35 次
- Generative Model-Based Testing on Decision-Making PoliciesZhuo Li, Xiongfei Wu, Derui Zhu, Mingfei Cheng 等ASE 2023 · 被引用 14 次
- Metamorphic relations via relaxations: an approach to obtain oracles for action-policy testingHasan Ferit Eniser, Timo P. Gros, Valentin Wüstholz, Jörg Hoffmann 等ISSTA 2022 · 被引用 14 次
相关 Paper
- MG-Fuzz: Model-Guided Fuzzing for Unsafe Scenario Discovery in Autonomous Driving SystemsYulong Lyu, Ruiqi Hong, Jiawan Wang, Jun Sun 等ISSTA 2026
- MoFuzz: A Fuzzer Suite for Testing Model-Driven Software Engineering ToolsHoang Lam Nguyen, Nebras Nassar, Timo Kehrer, Lars GrunskeASE 2020 · 被引用 12 次
- Curiosity-Driven Testing for Sequential Decision-Making ProcessJunda He, Zhou Yang, Jieke Shi, Chengran Yang 等ICSE 2024 · 被引用 6 次
- Reinforcement Learning-based Hierarchical Seed Scheduling for Greybox FuzzingJinghan Wang, Chengyu Song, Heng YinNDSS 2021
- On Interaction Effects in Greybox FuzzingKonstantinos Kitsios, Marcel Böhme, Alberto BacchelliICSE 2026
