ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
Suyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee, YunSeok Choi, Jee-Hyong Lee
摘要
As Large Language Models (LLMs) have become capable of generating long and descriptive code summaries, accurate and reliable evaluation of factual consistency has become a critical challenge. However, previous evaluation methods are primarily designed for short summaries of isolated code snippets. Consequently, they struggle to provide fine-grained evaluation of multi-sentence functionalities and fail to accurately assess dependency context commonly found in real-world code summaries. To address this, we propose ReFEree, a referencefree and fine-grained method for evaluating factual consistency in real-world code summaries. We define factual inconsistency criteria specific to code summaries and evaluate them at the segment level using these criteria along with dependency information. These segment-level results are then aggregated into a fine-grained score. We construct a code summarization benchmark with human-annotated factual consistency labels. The evaluation results demonstrate that ReFEree achieves the highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art. Our code and data are available at https: //github.com/bsy99615/ReFEree.git . Introduction Recent advances in Large Language Models (LLMs), such as GPT-4, have made it feasible to automatically generate long and descriptive code summaries (Achiam et al., 2023; Sun et al., 2024) . LLM-powered assistants such as OpenAI's Codex (Chen et al., 2021), GitHub Copilot (GitHub, 2021), and Anthropic's Claude-Code (Anthropic, 2025) are increasingly integrated into real-world development workflows to assist engineers in understanding and reviewing code. However, when the generated summary does not accurately reflect the code's actual implementation,
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 被引用 317 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
相关 Paper
- Do Automatic Factuality Metrics Measure Factuality? A Critical EvaluationSanjana Ramprasad, Byron C. WallaceNeurIPS 2025 · 被引用 13 次
- Calibration of Large Language Models on Code SummarizationYuvraj Virk, Premkumar T. Devanbu, Toufique AhmedFSE 2025 · 被引用 5 次
- FIZZ: Factual Inconsistency Detection by Zoom-in Summary and Zoom-out DocumentJoonho Yang, Seunghyun Yoon, Byeongjeong Kim, Hwanhee LeeEMNLP 2024 · 被引用 3 次
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri 等EMNLP 2023 · 被引用 29 次
- Source Code Summarization in the Era of Large Language ModelsWeisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 等ICSE 2025 · 被引用 37 次
