Refining and Reusing Annotation Guidelines for LLM Annotation
Kon Woo Kim, Jin-Dong Kim, Akiko Aizawa
Abstract
While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized conventions of gold-standard benchmarks. We propose the systematic reuse and refinement of annotation guidelines as an alignment mechanism, introducing an iterative moderation framework that simulates the early phases of annotation projects. We evaluate three hypotheses: (1) the efficacy of guideline integration, (2) the advantage of reasoning optimized models, and (3) the viability of moderation under minimal supervision. Testing across biomedical NER tasks (NCBI Disease, BC5CDR, BioRED) with three LLM families (GPT, Gemini, DeepSeek), our results empirically confirm all three hypotheses. While the iterative moderation framework shows good potential in effectively refining guidelines, our analysis also reveals substantial room for improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Is GPT-3 a Good Data Annotator?Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia et al.ACL 2023 · 133 citations
- GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity RecognitionShizhou Huang, Bo Xu, Yang Yu, Changqun Li et al.AAAI 2025 · 1 citation
- Named Entity Recognition with Small Strongly Labeled and Large Weakly Labeled DataHaoming Jiang, Danqing Zhang, Tianyu Cao, Bing Yin et al.ACL 2021
Related papers
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
- DiZiNER: Disagreement-guided Instruction Refinement via Simulating Pilot Annotation for Zero-shot Named Entity RecognitionSiun Kim, Hyung-Jin YoonACL 2026
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
- Generating Novel Leads for Drug Discovery Using LLMs with Logical FeedbackShreyas Bhat Brahmavar, Ashwin Srinivasan, Tirtharaj Dash, Sowmya Ramaswamy Krishnan et al.AAAI 2024 · 23 citations
- Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in MedicineMaxime Griot, Jean Vanderdonckt, Demet Yüksel, Coralie HemptinneACL 2025
