NIODebugger: A Novel Approach to Repair Non-Idempotent-Outcome Tests with LLM-Based Agent
Kaiyao Ke
Abstract
Flaky tests, characterized by inconsistent results across repeated executions, present significant challenges in software testing, especially during regression testing. Recently, there has been emerging research interest in non-idempotentoutcome (NIO) flaky tests-tests that pass on the initial run but fail on subsequent executions within the same environment. Despite progress in utilizing Large Language Models (LLMs) to address flaky tests, existing methods have not tackled NIO flaky tests. The limited context window of LLMs restricts their ability to incorporate relevant source code beyond the test method itself, often overlooking crucial information needed to address state pollution, which is the root cause of NIO flakiness. This paper introduces NIODebugger, the first framework to utilize an LLM-based agent to repair flaky tests. NIODebugger features a three-phase design: detection, exploration, and fixing. In the detection phase, dynamic analysis collects stack traces and custom test execution logs from multiple test runs, which helps in understanding accumulative state pollution. During the exploration phase, the LLM-based agent provides instructions for extracting relevant source code associated with test flakiness. In the fixing phase, NIODebugger repairs the tests using the information gathered from the previous phases. NIODebugger can be integrated with multiple LLMs, achieving patching success rates ranging from 11.63% to 58.72%. Its best-performing variant, NIODebugger-GPT-4, successfully generated correct patches for 101 out of 172 previously unknown NIO tests across 20 largescale open-source projects. We submitted pull requests for all generated patches; 58 have been merged, only 1 was rejected, and the remaining 42 are pending. The Java implementation of NIODebugger is provided as a Maven plugin accessible at https://github.com/kaiyaok2/NIOInspector.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 310a3606-9e18-43dd-b354-2dc819f2d552Cited by top-tier papers2
- SeCuRepair: Semantics-Aligned, Curriculum-Driven, and Reasoning-Enhanced Vulnerability Repair FrameworkChengran Yang, Ting Zhang, Jinfeng Jiang, Xin Zhou et al.ACL 2026 · 2 citations
- Wired for Reuse: Automating Context-Aware Code Adaptation in IDEs via LLM-Based AgentTaiming Wang, Yanjie Jiang, Chunhao Dong, Yuxia Zhang et al.ASE 2025
Related papers
- Preempting Flaky Tests via Non-Idempotent-Outcome TestsAnjiang Wei, Pu Yi, Zhengxi Li, Tao Xie et al.ICSE 2022 · 20 citations
- FlakyGuard: Automatically Fixing Flaky Tests at Industry ScaleChengpeng Li, Farnaz Behrang, August Shi, Peng LiuASE 2025 · 1 citation
- UniDebugger: Hierarchical Multi-Agent Framework for Unified Software DebuggingCheryl Lee, Chunqiu Steven Xia, Longji Yang, Jen-tse Huang et al.EMNLP 2025
- RepairAgent: An Autonomous, LLM-Based Agent for Program RepairIslem Bouzenia, Premkumar T. Devanbu, Michael PradelICSE 2025 · 54 citations
- Dockerfile Flakiness: Characterization and RepairTaha Shabani, Noor Nashid, Parsa Alian, Ali MesbahICSE 2025 · 2 citations
