CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
Zhengmin Yu, Yuan Zhang, Ming Wen, Yinan Nie, Wenhui Zhang, Min Yang
Abstract
Project building is pivotal to support various program analysis tasks, such as generating intermediate representation code for static analysis and preparing binary code for vulnerability reproduction. However, automating the building process for C/C++ projects is a highly complex endeavor, involving tremendous technical challenges, such as intricate dependency management, diverse build systems, varied toolchains, and multifaceted error handling mechanisms. Consequently, building C/C++ projects often proves to be difficult in practice, hindering the progress of downstream applications. Unfortunately, research on facilitating the building of C/C++ projects remains to be inadequate. The emergence of Large Language Models (LLMs) offers promising solutions to automated software building. Trained on extensive corpora, LLMs can help unify diverse build systems through their comprehension capabilities and address complex errors by leveraging tacit knowledge storage. Moreover, LLM-based agents can be systematically designed to dynamically interact with the environment, effectively managing dynamic building issues. Motivated by these opportunities, we first conduct an empirical study to systematically analyze the current challenges in the C/C++ project building process. Particularly, we observe that most popular C/C++ projects encounter an average of five errors when relying solely on the default build systems. Based on our study, we develop an automated build system called CXXCrafter to specifically address the above-mentioned challenges, such as dependency resolution. Our evaluation on open-source software demonstrates that CXXCrafter achieves a success rate of 78% in project building. Specifically, among the Top100 dataset, 72 projects are built successfully by both CXXCrafter and manual efforts, 3 by CXXCrafter only, and 14 manually only. Despite the slightly lower performance, CXXCrafter can save tremendous manual efforts and can also be easily applied to a wider range of applications automatically. CCS Concepts: • Software and its engineering → Software notations and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9b99ba8-0fa0-4c1f-93da-325e443178a7Cited by top-tier papers6
- Diffploit: Facilitating Cross-Version Exploit Migration for Open Source Library VulnerabilitiesZirui Chen, Zhipeng Xue, Jiayuan Zhou, Xing Hu et al.ICSE 2026
- Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding StandardsZejun Zhang, Yixin Gan, Zhenchang Xing, Tian Zhang et al.FSE 2026
- Fact-Aligned and Template-Constrained Static Analyzer Rule Enhancement with LLMsZongze Jiang, Ming Wen, Ge Wen, Hai JinASE 2025
- Guarding the Lifeline: A First Look and Automated Defect Diagnosis for ROS Central IndexWeijie Sun, Huiyan Wang, Ying Wang, Chang XuISSTA 2026
- EvidenT: An Evidence-Preserving Framework for Iterative System-Level Package RepairChenyu Zhao, Minghua Ma, Shenglin Zhang, Zeshun Huang et al.ISSTA 2026
Builds on10
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- Large Language Models for Code Analysis: Do LLMs Really Do Their Job?Chongzhou Fang, Ning Miao, Shaurya Srivastav, Jialin Liu et al.USENIX Security 2024 · 110 citations
- Towards Understanding Third-party Library Dependency in C/C++ EcosystemWei Tang, Zhengzi Xu, Chengwei Liu, Jiahui Wu et al.ASE 2022 · 64 citations
- Understanding build issue resolution in practice: symptoms and fix patternsYiling Lou, Zhenpeng Chen, Yanbin Cao, Dan Hao et al.FSE 2020 · 39 citations
Related papers
- AutoBaxBuilder: Bootstrapping Code Security BenchmarkingTobias von Arx, Niels Mündler, Mark Vero, Maximilian Baader et al.ICML 2026 · 1 citation
- Interleaving Large Language Models for Compiler TestingYunbo Ni, Shaohua LiOOPSLA 2025 · 4 citations
- Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ BugsJian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu et al.ASE 2025
- CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent SystemLi Hu, Guoqiang Chen, Xiuwei Shang, Shaoyin Cheng et al.ACL 2025
- You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary ProjectsIslem Bouzenia, Michael PradelISSTA 2025 · 17 citations
