Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs during Code Adaptation
Tanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yuxin Zhao, Yao Lu, Zezhou Tang, Wenyu Xu, Longfei Sun, Changrong Xie, Kang Yang, Yue Yu
Abstract
Code adaptation is a fundamental but challenging task in software development, requiring developers to modify existing code for new contexts. A key challenge is to resolve Context Adaptation Bugs (CtxBugs) , which occurs when code correct in its original context violates constraints in the target environment. Unlike isolated bugs, CtxBugs cannot be resolved through local fixes and require cross-context reasoning to identify semantic mismatches. Overlooking them may lead to critical failures in adaptation. Although Large Language Models (LLMs) show great potential in automating code-related tasks, their ability to resolve CtxBugs remains a significant and unexplored obstacle to their practical use in code adaptation. To bridge this gap, we propose CtxBugGen , a novel framework for generating CtxBugs to evaluate LLMs. Its core idea is to leverage LLMs’ tendency to generate plausible but context-free code when contextual constraints are absent. The framework generates CtxBugs through a four-step process to ensure their relevance and validity: (1) Selection of four established context-aware adaptation tasks from the literature, (2) Perturbation via task-specific rules to induce CtxBugs from LLMs while ensuring their plausibility, (3) Generation of candidate variants by prompting LLMs without any context constraints and (4) Identification of valid CtxBugs through syntactic differencing and test execution in the target context. Based on the benchmark constructed by CtxBugGen , we conduct an empirical study with four state-of-the-art LLMs. Our results reveal their unsatisfactory performance in CtxBug resolution. The best performing LLM, Kimi-K2, achieves 55.93% on Pass@1 and resolves just 52.47% of CtxBugs . The presence of CtxBugs degrades LLMs’ adaptation performance by up to 30%. Failure analysis indicates that LLMs often overlook CtxBugs and replicate them in their outputs. This suggests that LLMs overly focus on the local code correctness of the reused code while ignoring its compatibility in the target context. Our study highlights a critical weakness in LLMs’ cross-context reasoning and emphasize the need for new methods to enhance their context awareness for reliable code adaptation. The replication package for this paper is at https://github.com/ztwater/CtxBugGen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
- Stack Overflow Considered Harmful? The Impact of Copy&Paste on Android Application SecurityFelix Fischer, Konstantin Böttinger, Huang Xiao, Christian Stransky et al.S&P 2017 · 293 citations
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 223 citations
Related papers
- Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringTanghaoran Zhang, Yue Yu, Xinjun Mao, Shangwen Wang et al.ICSE 2025 · 3 citations
- Large Language Models of Code Fail at Completing Code with Potential BugsTuan Dinh, Jinman Zhao, Samson Tan, Renato Negrinho et al.NeurIPS 2023 · 59 citations
- AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet AdaptationTanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yuxin Zhao et al.ASE 2025 · 1 citation
- Kimi-Dev: Agentless Training as Skill Prior for SWE-agentsZonghan Yang, Shengjie Wang, Kelin Fu, Wenyang He et al.ICLR 2026 · 34 citations
- LLM-Powered Test Case Generation for Detecting Bugs in Plausible ProgramsKaibo Liu, Zhenpeng Chen, Yiyang Liu, Jie M. Zhang et al.ACL 2025 · 20 citations
