Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain
Marcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar, Gail E. Kaiser, Suman Jana, Baishakhi Ray
摘要
Code Large Language Models (Code LLMs) are being increasingly employed in real-life applications, so evaluating them is critical. While the conventional accuracy evaluates the performance of Code LLMs on a set of individual tasks, their self-consistency across different tasks is overlooked. Intuitively, a trustworthy model should be self-consistent when generating natural language specifications for its own code and generating code for its own specifications. Failure to preserve self-consistency reveals a lack of understanding of the shared semantics underlying natural language and programming language, and therefore undermines the trustworthiness of a model. In this paper, we first formally define the self-consistency of Code LLMs and then design a framework, IdentityChain, which effectively and efficiently evaluates the self-consistency and conventional accuracy of a model at the same time. We study eleven Code LLMs and show that they fail to preserve self-consistency, which is indeed a distinct aspect from conventional accuracy. Furthermore, we show that IdentityChain can be used as a model debugging tool to expose weaknesses of Code LLMs by demonstrating three major weaknesses that we identify in current models using IdentityChain. Our code is available at https://github.com/marcusm117/IdentityChain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Unsupervised Evaluation of Code LLMs with Round-Trip CorrectnessMiltiadis Allamanis, Sheena Panthaplackel, Pengcheng YinICML 2024 · 被引用 26 次
- Automated Program Refinement: Guide and Verify Code Large Language Model with Refinement CalculusYufan Cai, Zhe Hou, David Sanán, Xiaokun Luan 等POPL 2025 · 被引用 20 次
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM HallucinationsBuyun Liang, Liangzu Peng, Jinqi Luo, Darshan Thaker 等NeurIPS 2025 · 被引用 11 次
- ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented GenerationPengcheng Huang, Zhenghao Liu, Yukun Yan, Haiyan Zhao 等NeurIPS 2025 · 被引用 11 次
- tnGPS: Discovering Unknown Tensor Network Structure Search Algorithms via Large Language Models (LLMs)Junhua Zeng, Chao Li, Zhun Sun, Qibin Zhao 等ICML 2024 · 被引用 10 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
相关 Paper
- CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modulesHung Le, Hailin Chen, Amrita Saha, Akash Gokul 等ICLR 2024 · 被引用 73 次
- Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language ModelsYanlin Wang, Tianyue Jiang, Mingwei Liu, Jiachi Chen 等FSE 2025 · 被引用 6 次
- TACO: Trust Assessment of Large Language Models in Coding Assistance TasksShihao Weng, Yang Feng, Jincheng Li, Yining Yin 等ICSE 2026
- EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence CheckingAnjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen 等EMNLP 2025
- CodeJudge: Evaluating Code Generation with Large Language ModelsWeixi Tong, Tianyi ZhangEMNLP 2024 · 被引用 25 次
