TACO: Trust Assessment of Large Language Models in Coding Assistance Tasks
Shihao Weng, Yang Feng, Jincheng Li, Yining Yin, Zhenlun Zhang, Lyuxi Liu, Jia Liu
摘要
Large Language Models (LLMs) have rapidly become integral to software development workflows, particularly in coding assistance tasks (CAT) such as debugging, implementation, and code optimization. However, the trustworthiness of LLM-generated responses remains a critical concern, as hallucinations (incorrect or misleading outputs) can severely hinder developer productivity and software reliability. In this paper, we introduce TACO, a comprehensive framework for trust assessment of LLMs in CAT scenarios. TACO jointly evaluates both the code quality and the alignment with user intent of LLM responses, enabling fine-grained and interpretable trust evaluation. We construct two new benchmarks: TACO-Judge, a human-annotated dataset for validating evaluation methods, and TACO-Eval, a large-scale benchmark for assessing LLM performance on real-world CAT problems. Through extensive experiments, we (1) demonstrate the accuracy of TACO as an automated evaluator, (2) investigate whether LLMs exhibit self-preference when scoring their own outputs, and (3) compare the trustworthiness of several state-of-the-art LLMs in CAT scenarios. Our results highlight both the promise and limitations of current LLMs, and establish TACO as a reliable tool for their evaluation in software engineering contexts.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao 等AAAI 2025 · 被引用 41 次
- ETF: An Entity Tracing Framework for Hallucination Detection in Code SummariesKishan Maharaj, Vitobha Munigala, Srikanth G. Tamilselvam, Prince Kumar 等ACL 2025 · 被引用 4 次
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen 等FSE 2025 · 被引用 1 次
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma 等FSE 2024 · 被引用 8 次
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChainMarcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar 等ICLR 2024 · 被引用 39 次
