Environment-Aware Code Generation: How far are We?
Tongtong Wu, Rongyi Chen, Wenjie Du, Suyu Ma, Guilin Qi, Zhenchang Xing, Shahram Khadivi, Ramesh Periyathambi, Gholamreza Haffari
Abstract
Recent progress in large language models (LLMs) has led to impressive code generation capabilities. However, existing evaluations of LLMs primarily focus on generating isolated, small-scale code units (e.g., single functions or statements) under default or unspecified software environments. As a result, it remains unclear whether LLMs can reliably generate executable code tailored to specific user environments. To fill this knowledge gap, we make the first systematic study of Environment-Aware Code Generation (EACG), which requires generating code that is both functionally correct and directly executable under arbitrary software configurations. To support realistic evaluation, we introduce VersiBCB, a benchmark featuring multi-package, executable-verified, and deprecation-aware, reflecting complex and evolving software environments that are often overlooked in prior datasets. Building on this benchmark, we explore three orthogonal adaptation axes: data, parameters, and cache, and further develop representative strategies for each. Our results reveal that existing LLMs struggle with environment-specific code generation, but our adaptation strategies yield improvements in environment compatibility and executability. These findings highlight critical challenges and opportunities for deploying LLMs in practical software engineering workflows.
• Software and its engineering → Software development techniques; • Computing methodologies → Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bbbf3a3-b86f-428a-93a7-1a71afd79443Cited by top-tier papers1
Ask how each one uses itBuilds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- Evaluating Large Language Models in Class-Level Code GenerationXueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang et al.ICSE 2024 · 118 citations
Related papers
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang et al.FSE 2026
- HumanEvo: An Evolution-Aware Benchmark for More Realistic Evaluation of Repository-Level Code GenerationDewu Zheng, Yanlin Wang, Ensheng Shi, Ruikai Zhang et al.ICSE 2025 · 2 citations
- Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'Shanchao Liang, Nan Jiang, Yiran Hu, Lin TanACL 2025 · 9 citations
- CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and GenerationWeixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li et al.ACL 2024
- When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task DescriptionsMaya Larbi, Amal Akli, Mike Papadakis, Rihab Bouyousfi et al.ICSE 2026
