Detecting Errors in AI-Generated Annotations: When and Why Semantic Neighbors Help
Na Di, Ling Li, Zhe Tang, Hao Cheng, Jinlong Pang, Jiaheng Wei, Zhaowei Zhu
Abstract
Large language models (LLMs) and vision-language models (VLMs) have emerged as efficient annotators for tasks such as generation and classification. While these models offer significant cost and speed advantages over human annotation, a critical challenge remains: existing self-evaluation methods, such as LLM-as-judge, often lack reliable reference-based calibration for error detection. We address this limitation by introducing SAGE (Semantic-Anchored JudGmEnt), a method that leverages semantically similar samples retrieved via -nearest-neighbor as references for annotation verification. We provide a theoretical framework that derives a closed-form expression for the error detection AUROC, which can be decomposed into three factors: intrinsic separability, reference-induced mean shift, and noise reduction through averaging. This decomposition reveals when semantic neighbors help (when references are both semantically matched and correct) and why (by providing reference-based calibration that raises scores for correct annotations and lowers scores for incorrect ones). Experiments on LLM generation, VLM captioning, and classification tasks validate our theoretical framework: SAGE improves error detection when semantic neighbors provide reliable reference-based calibration, and our decomposition offers insights into when direct scoring or alternative strategies may be preferred. Our code is available at https://github.com/dina-1205/SAGE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d5c1ade-fe98-46e6-9cf3-06d3efcc05b6Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 1 citation
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li et al.AAAI 2026
- Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the dataFlorian E. Dorner, Vivian Yvonne Nastl, Moritz HardtICLR 2025
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge VerificationKanghoon Yoon, Minsub Kim, Sungjae Lee, Joonhyung Lee et al.ICML 2026 · 4 citations
- Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceOmer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor et al.EMNLP 2025 · 9 citations
