ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
Hojae Han, Jaejin Kim, Jaeseok Yoo, Youngwon Lee, Seung-won Hwang
Abstract
This paper aims to extend the code generation capability of large language models (LLMs) to automatically manage comprehensive software requirements from given textual descriptions. Such requirements include both functional (i.e. achieving expected behavior for inputs) and non-functional (e.g., time/space performance, robustness, maintainability) requirements. However, textual descriptions can either express requirements verbosely or may even omit some of them. We introduce ARCHCODE, a novel framework that leverages in-context learning to organize requirements observed in descriptions and to extrapolate unexpressed requirements from them. ARCHCODE generates requirements from given descriptions, conditioning them to produce code snippets and test cases. Each test case is tailored to one of the requirements, allowing for the ranking of code snippets based on the compliance of their execution results with the requirements. Public benchmarks show that ARCHCODE enhances to satisfy functional requirements, significantly improving Pass@k scores. Furthermore, we introduce HumanEval-NFR, the first evaluation of LLMs' non-functional requirements in code generation, demonstrating ARCHCODE's superiority over baseline methods. The implementation of ARCHCODE and the HumanEval-NFR benchmark are both publicly accessible. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdabbf18-449c-49a0-aff2-cdbd46805b32Cited by top-tier papers5
- GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment DesignWen-Fan Wang, Ting-Ying Lee, Chien-Ting Lu, Che-Wei Hsu et al.UIST 2025 · 4 citations
- Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph ReasoningXingrui Zhuo, Jiapu Wang, Gongqing Wu, Zhongyuan Wang et al.ICLR 2026 · 2 citations
- Intention Chain-of-Thought Prompting with Dynamic Routing for Code GenerationShen Li, Li Huang, Shaoxiong Zhan, Weifeng Sun et al.AAAI 2026 · 1 citation
- TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging of LLM-Generated CodeJiangping Huang, Wenguang Ye, Weisong Sun, Jian Zhang et al.ICSE 2026 · 1 citation
- GiFT: Gibbs Fine-Tuning for Code GenerationHaochen Li, Wanjin Feng, Xin Zhou, Zhiqi ShenACL 2025
Builds on10
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov et al.ICML 2023 · 318 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
Related papers
- When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task DescriptionsMaya Larbi, Amal Akli, Mike Papadakis, Rihab Bouyousfi et al.ICSE 2026
- Compiling Large Multi-modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven PerspectiveWeiyu Kong, Yun Lin, Xiwen Teoh, Duc-Minh Nguyen et al.ISSTA 2026 · 1 citation
- The First Prompt Counts the Most! An Evaluation of Large Language Models on Iterative Example-Based Code GenerationYingjie Fu, Bozhou Li, Linyi Li, Wentao Zhang et al.ISSTA 2025 · 3 citations
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu et al.ACL 2023 · 104 citations
- RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development PracticesJia Li, Hongyi Deng, Yiran Zhang, Kechi Zhang et al.FSE 2026
