ETF: An Entity Tracing Framework for Hallucination Detection in Code Summaries
Kishan Maharaj, Vitobha Munigala, Srikanth G. Tamilselvam, Prince Kumar, Sayandeep Sen, Palani Kodeswaran, Abhijit Mishra, Pushpak Bhattacharyya
摘要
Recent advancements in large language models (LLMs) have significantly enhanced their ability to understand both natural language and code, driving their use in tasks like natural language-to-code (NL2Code) and code summarisation. However, LLMs are prone to hallucination, outputs that stray from intended meanings. Detecting hallucinations in code summarisation is especially difficult due to the complex interplay between programming and natural languages. We introduce a first-of-its-kind dataset, CodeSumEval, with 10K samples, curated specifically for hallucination detection in code summarisation. We further propose a novel Entity Tracing Framework (ETF) that a) utilises static program analysis to identify code entities from the program and b) uses LLMs to map and verify these entities and their intents within generated code summaries. Our experimental analysis demonstrates the framework's effectiveness, leading to a 73% F1 score. The proposed approach provides a method for detecting hallucinations by tracing entities from the summary to the code, allowing us to evaluate summary accuracy and localise the error within the summary.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code SummarizationSuyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee 等ACL 2026 · 被引用 1 次
- AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering AgentsZhengran Zeng, Yixin Li, Rui Xie, Wei Ye 等ISSTA 2026
它引用的顶会 Paper5
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- Ask Me Anything: A simple strategy for prompting language modelsSimran Arora, Avanika Narayan, Mayee F. Chen, Laurel J. Orr 等ICLR 2023 · 被引用 74 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
- Code and Named Entity Recognition in StackOverflowJeniya Tabassum, Mounica Maddela, Wei Xu, Alan RitterACL 2020 · 被引用 9 次
- We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMsJoseph Spracklen, Raveen Wijewickrama, A. H. M. Nazmus Sakib, Anindya Maiti 等USENIX Security 2025
相关 Paper
- Hallucinations in LLM-Based Code Summarization: Unveiling, Detection, and MitigationGuanghua Wan, Yuanning Feng, Yao Wan, Zhaoyang Chu 等FSE 2026
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao 等AAAI 2025 · 被引用 41 次
- TACO: Trust Assessment of Large Language Models in Coding Assistance TasksShihao Weng, Yang Feng, Jincheng Li, Yining Yin 等ICSE 2026
- Contrastive Error Attribution for Finetuned Language ModelsFaisal Ladhak, Esin Durmus, Tatsunori HashimotoACL 2023
- The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language ModelsJunyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng 等ACL 2024 · 被引用 49 次
