Causality-Aided Evaluation and Explanation of Large Language Model-Based Code Generation
Zhenlan Ji, Pingchuan Ma, Zongjie Li, Zhaoyu Wang, Shuai Wang
Abstract
While code generation has been widely used in various software development scenarios, the quality of the generated code is not guaranteed. This has been a particular concern in the era of large language models (LLM)-based code generation, where LLMs, deemed a complex and powerful black-box model, are instructed by a high-level natural language specification, namely a prompt, to generate code. Nevertheless, effectively evaluating and explaining the code generation capability of LLMs is inherently challenging, given the complexity of LLMs and the lack of transparency. Inspired by recent progress in causality analysis and its software engineering applications, this paper proposes a causality-driven approach to systematically analyze prompt-code causal relationships. However, this endeavor faces three key technical challenges: (1) representing textual prompts and code in a canonical form, (2) establishing causal relations between high-level concepts and code features, and (3) systematically analyzing diverse prompt variations. To address these challenges, we first propose a novel causal graph-based representation of the prompt and the generated code, which is established over the fine-grained, human-understandable concepts in the input prompts. The formed causal graph is then used to identify the causal relations between the prompt and the derived code. We illustrate the insights that our framework can provide by studying over four popular LLMs with over 12 prompt adjustment strategies. The results of these studies illustrate the potential of our technique to provide insights into LLM effectiveness and aid end-users in understanding predictions. Additionally, we demonstrate that our approach provides actionable insights to improve the quality of the LLM-generated code by properly calibrating the prompt.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3539e3e9-5b1a-4bdc-8df8-c79e4a32fbc2Cited by top-tier papers2
- Don’t Use a Cannon to Kill a Fly: Lightweight Model Editing for LLMs to Correct Deprecated API RecommendationsGuancheng Lin, Xiao Yu, Jacky Keung, Xing Hu et al.ISSTA 2026
- CAM: A Causality-Based Analysis Framework for Multi-agent Code Generation SystemsZongyi Lyu, Zhenlan Ji, Songqiang Chen, Liwen Wang et al.ISSTA 2026
Related papers
- Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringTanghaoran Zhang, Yue Yu, Xinjun Mao, Shangwen Wang et al.ICSE 2025 · 3 citations
- Causal Order: The Key to Leveraging Imperfect Experts in Causal InferenceAniket Vashishtha, Abbavaram Gowtham Reddy, Abhinav Kumar, Saketh Bachu et al.ICLR 2025
- Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit TestsJunda Zhao, Shurui Zhou, Eldan CohenISSTA 2026 · 1 citation
- A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated CodeAlejandro Velasco, Daniel Rodriguez-Cardenas, Dipin Khati, David N. Palacio et al.ICSE 2026
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph InferenceBen Finkelshtein, Silviu Cucerzan, Sujay Kumar Jauhar, Ryen W WhiteICLR 2026 · 5 citations
