Refining and Reusing Annotation Guidelines for LLM Annotation
Kon Woo Kim, Jin-Dong Kim, Akiko Aizawa
摘要
While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized conventions of gold-standard benchmarks. We propose the systematic reuse and refinement of annotation guidelines as an alignment mechanism, introducing an iterative moderation framework that simulates the early phases of annotation projects. We evaluate three hypotheses: (1) the efficacy of guideline integration, (2) the advantage of reasoning optimized models, and (3) the viability of moderation under minimal supervision. Testing across biomedical NER tasks (NCBI Disease, BC5CDR, BioRED) with three LLM families (GPT, Gemini, DeepSeek), our results empirically confirm all three hypotheses. While the iterative moderation framework shows good potential in effectively refining guidelines, our analysis also reveals substantial room for improvement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Is GPT-3 a Good Data Annotator?Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia 等ACL 2023 · 被引用 133 次
- GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity RecognitionShizhou Huang, Bo Xu, Yang Yu, Changqun Li 等AAAI 2025 · 被引用 1 次
- Named Entity Recognition with Small Strongly Labeled and Large Weakly Labeled DataHaoming Jiang, Danqing Zhang, Tianyu Cao, Bing Yin 等ACL 2021
相关 Paper
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon 等ICLR 2026 · 被引用 13 次
- DiZiNER: Disagreement-guided Instruction Refinement via Simulating Pilot Annotation for Zero-shot Named Entity RecognitionSiun Kim, Hyung-Jin YoonACL 2026
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing 等AAAI 2026 · 被引用 4 次
- Generating Novel Leads for Drug Discovery Using LLMs with Logical FeedbackShreyas Bhat Brahmavar, Ashwin Srinivasan, Tirtharaj Dash, Sowmya Ramaswamy Krishnan 等AAAI 2024 · 被引用 23 次
- Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in MedicineMaxime Griot, Jean Vanderdonckt, Demet Yüksel, Coralie HemptinneACL 2025
