Comfrey: Mitigating Integration Failures in LLM-enabled Software at Run-Time
Yuchen Shao, Yuheng Huang, Jiazhen Zou, Yuling Shi, Long Yang, Lei Ma, Ting Su, Chengcheng Wan
Abstract
Due to the unrestricted outputs of LLMs and strict requirements of software components, integration failures are widespread in software that incorporates LLM agents and retrieval-augmented generation (RAG). Even seemingly correct LLM/RAG responses can trigger software misbehaviors if they violate these requirements.
In this paper, we conduct an empirical study to understand integration failures in real-world LLM-enabled applications. Guided by this study, we present Comfrey [1], a runtime framework that adapts the LLM agent and RAG responses to meet software requirements, serving as a middle layer between AI and software components. It automatically detects and resolves potential integration failures through a three-stage workflow, ensuring component compatibility. Our evaluation with a variety of open-source applications demonstrates that Comfrey detects 75.1% and prevents 63.3% of potential integration failures with 8.4% overhead, significantly outperforming the baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cef5006a-700f-4d44-868b-2e6192a79157Builds on50
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Are LLMs Correctly Integrated into Software Systems?Yuchen Shao, Yuheng Huang, Jiawei Shen, Lei Ma et al.ICSE 2025 · 4 citations
- Run-Time Prevention of Software Integration Failures of Machine Learning APIsChengcheng Wan, Yuhan Liu, Kuntai Du, Henry Hoffmann et al.OOPSLA 2023 · 5 citations
- Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs during Code AdaptationTanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yuxin Zhao et al.FSE 2026
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 95 citations
- On Automating Configuration Dependency Validation via Retrieval-Augmented GenerationSebastian Simon, Alina Mailach, Johannes Dorn, Norbert SiegmundASE 2025
