TACO: Trust Assessment of Large Language Models in Coding Assistance Tasks
Shihao Weng, Yang Feng, Jincheng Li, Yining Yin, Zhenlun Zhang, Lyuxi Liu, Jia Liu
Abstract
Large Language Models (LLMs) have rapidly become integral to software development workflows, particularly in coding assistance tasks (CAT) such as debugging, implementation, and code optimization. However, the trustworthiness of LLM-generated responses remains a critical concern, as hallucinations (incorrect or misleading outputs) can severely hinder developer productivity and software reliability. In this paper, we introduce TACO, a comprehensive framework for trust assessment of LLMs in CAT scenarios. TACO jointly evaluates both the code quality and the alignment with user intent of LLM responses, enabling fine-grained and interpretable trust evaluation. We construct two new benchmarks: TACO-Judge, a human-annotated dataset for validating evaluation methods, and TACO-Eval, a large-scale benchmark for assessing LLM performance on real-world CAT problems. Through extensive experiments, we (1) demonstrate the accuracy of TACO as an automated evaluator, (2) investigate whether LLMs exhibit self-preference when scoring their own outputs, and (3) compare the trustworthiness of several state-of-the-art LLMs in CAT scenarios. Our results highlight both the promise and limitations of current LLMs, and establish TACO as a reliable tool for their evaluation in software engineering contexts.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f531cd60-2e58-4af9-bf9a-9349ff07e208Related papers
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao et al.AAAI 2025 · 41 citations
- ETF: An Entity Tracing Framework for Hallucination Detection in Code SummariesKishan Maharaj, Vitobha Munigala, Srikanth G. Tamilselvam, Prince Kumar et al.ACL 2025 · 4 citations
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen et al.FSE 2025 · 1 citation
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma et al.FSE 2024 · 8 citations
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChainMarcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar et al.ICLR 2024 · 39 citations
