Can Large Language Models Write Parallel Code?
Daniel Nichols, Joshua Hoke Davis, Zhaojun Xie, Arjun Rajaram, Abhinav Bhatele
Abstract
Large language models are increasingly becoming a popular tool for software development. Their ability to model and generate source code has been demonstrated in a variety of contexts, including code completion, summarization, translation, and lookup. However, they often struggle to generate code for complex programs. In this paper, we study the capabilities of state-of-the-art language models to generate parallel code. In order to evaluate language models,we create a benchmark, ParEval, consisting of prompts that represent 420 different coding tasks related to scientific and parallel computing. We use ParEval to evaluate the effectiveness of several state-of-the-art open- and closed-source language models on these tasks. We introduce novel metrics for evaluating the performance of generated code, and use them to explore how well each large language model performs for 12 different computational problem types and six different parallel programming models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52d024ff-3449-405a-b9d5-63ff9f120d9eCited by top-tier papers10
- CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel ProgrammingAli TehraniJamsaz, Arijit Bhattacharjee, Le Chen, Nesreen K. Ahmed et al.NeurIPS 2024 · 36 citations
- Generative AI Uses and Risks for Knowledge Workers in a Science OrganizationKelly B. Wagman, Matthew T. Dearing, Marshini ChettyCHI 2025 · 16 citations
- From Large to Small: Transferring CUDA Optimization Expertise via Reasoning GraphJunfeng Gong, Zhiyi Wei, Junying Chen, Cheng Liu et al.ICLR 2026 · 10 citations
- Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code TranslationLe Chen, Nuo Xu, Winson Chen, Bin Lei et al.ACL 2026 · 6 citations
- ParaCodex: A Profiling-Guided Autonomous Coding Agent for Reliable Parallel Code Generation and TranslationErel Kaplan, Tomer Bitan, Lian Ghrayeb, Le Chen et al.ACL 2026 · 1 citation
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
Related papers
- McEval: Massively Multilingual Code EvaluationLinzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin et al.ICLR 2025 · 1 citation
- XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalMohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang et al.ACL 2024 · 21 citations
- CONCUR: Benchmarking LLMs for Concurrent Code GenerationJue Huang, Tarek Mahmud, Corina S. Pasareanu, Guowei YangISSTA 2026
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong et al.ASE 2024 · 7 citations
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu et al.ACL 2023 · 104 citations
