Lune

ICSE2026Top-tier venue

TACO: Trust Assessment of Large Language Models in Coding Assistance Tasks

Shihao Weng, Yang Feng, Jincheng Li, Yining Yin, Zhenlun Zhang, Lyuxi Liu, Jia Liu

2026Year

Abstract

Large Language Models (LLMs) have rapidly become integral to software development workflows, particularly in coding assistance tasks (CAT) such as debugging, implementation, and code optimization. However, the trustworthiness of LLM-generated responses remains a critical concern, as hallucinations (incorrect or misleading outputs) can severely hinder developer productivity and software reliability. In this paper, we introduce TACO, a comprehensive framework for trust assessment of LLMs in CAT scenarios. TACO jointly evaluates both the code quality and the alignment with user intent of LLM responses, enabling fine-grained and interpretable trust evaluation. We construct two new benchmarks: TACO-Judge, a human-annotated dataset for validating evaluation methods, and TACO-Eval, a large-scale benchmark for assessing LLM performance on real-world CAT problems. Through extensive experiments, we (1) demonstrate the accuracy of TACO as an automated evaluator, (2) investigate whether LLMs exhibit self-preference when scoring their own outputs, and (3) compare the trustworthiness of several state-of-the-art LLMs in CAT scenarios. Our results highlight both the promise and limitations of current LLMs, and establish TACO as a reliable tool for their evaluation in software engineering contexts.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get f531cd60-2e58-4af9-bf9a-9349ff07e208

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines