EvidenT: An Evidence-Preserving Framework for Iterative System-Level Package Repair
Chenyu Zhao, Minghua Ma, Shenglin Zhang, Zeshun Huang, Yongqian Sun, Chetan Bansal, Saravan Rajmohan, Dan Pei
摘要
Frequent toolchain updates and the expanding diversity of instruction set architectures (ISAs) have made large-scale system-level software package repair a critical task. Diagnosing and repairing build failures remains challenging due to heterogeneous failure evidence, complex dependency constraints, and architecture-specific build conventions. While recent LLM-based repair methods have shown promise for project-level source code fixes, they struggle with system-level repair where failures involve multi-language artifacts (e.g., build recipes, scripts, and source archives) and require iterative validation through external build services. In this paper, we first conduct a systematic empirical study of real-world system-level build failures. Our findings reveal that 72% of successful repairs primarily involve adjustments to build configurations, dependencies, or environment settings rather than isolated source-code modifications, suggesting that effective repair must prioritize packaging logic and iterative feedback. Motivated by these insights, we propose EvidenT, an evidence-preserving repair framework that decouples iteration-aware evidence management from tool execution. EvidenT comprises (1) an external Build Service for reproducible build execution and feedback; (2) an Evidence-Preserving Repair Controller that performs cross-modal fusion of repair history, knowledge context, and build artifacts; and (3) an automated Repair Orchestrator that executes a suite of modular tools for failure localization and system-level repair actions within a closed-loop validation environment. We evaluate EvidenT on a benchmark of 219 real-world RISC-V package build failures. EvidenT successfully repairs 118 packages (53.88%), substantially outperforming state-of-the-art agentic baselines (20.55%) and direct LLM-based repair (1.83%). To demonstrate its architectural generality, we extend EvidenT to other ISAs by updating only ISA-specific knowledge context. In preliminary experiments, it achieves success rates of 41.77% on aarch64 and 46.99% on x86_64, showcasing its robustness across diverse hardware ecosystems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 等EuroSys 2024 · 被引用 175 次
- RepairAgent: An Autonomous, LLM-Based Agent for Program RepairIslem Bouzenia, Premkumar T. Devanbu, Michael PradelICSE 2025 · 被引用 54 次
- MULAN: Multi-modal Causal Structure Learning and Root Cause Analysis for Microservice SystemsLecheng Zheng, Zhengzhang Chen, Jingrui He, Haifeng ChenWWW 2024 · 被引用 53 次
- Escaping dependency hell: finding build dependency errors with the unified dependency graphGang Fan, Chengpeng Wang, Rongxin Wu, Xiao Xiao 等ISSTA 2020 · 被引用 37 次
相关 Paper
- Root Cause Analysis of RISC-V Build Failures via LLM and MCTS ReasoningWeipeng Shuai, Jie Liu, Zhirou Ma, Liangyi Kang 等ASE 2025
- ExpeRepair: Dual-Memory Enhanced LLM-Based Repository-Level Program RepairFangwen Mu, Junjie Wang, Lin Shi, Song Wang 等FSE 2026 · 被引用 2 次
- PATCHAGENT: A Practical Program Repair Agent Mimicking Human ExpertiseZheng Yu, Ziyi Guo, Yuhang Wu, Jiahao Yu 等USENIX Security 2025
- LLM-based Agents for Automated Bug Fixing: How Far Are We?Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao PengICSE 2026
- Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesIslem Bouzenia, Michael PradelASE 2025 · 被引用 3 次
