GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction
Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle, German Rigau, Eneko Agirre
Abstract
Large Language Models (LLMs) combined with instruction tuning have made significant progress when generalizing to unseen tasks. However, they have been less successful in Information Extraction (IE), lagging behind task-specific models. Typically, IE tasks are characterized by complex annotation guidelines that describe the task and give examples to humans. Previous attempts to leverage such information have failed, even with the largest models, as they are not able to follow the guidelines out of the box. In this paper, we propose GoLLIE (Guidelinefollowing Large Language Model for IE), a model able to improve zero-shot results on unseen IE tasks by virtue of being fine-tuned to comply with annotation guidelines. Comprehensive evaluation empirically demonstrates that GoLLIE is able to generalize to and follow unseen guidelines, outperforming previous attempts at zero-shot information extraction. The ablation study shows that detailed guidelines are key for good results. Code, data, and models are publicly available: https://github.com/hitz-zentroa/GoLLIE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 007ffa3c-543f-4e43-bbb1-83cb5c856672Cited by top-tier papers31
- NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated DataSergei Bogdanov, Alexandre Constantin, Timothée Bernard, Benoît Crabbé et al.EMNLP 2024 · 29 citations
- PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity RecognitionJinghui Lu, Yanjie Wang, Ziwei Yang, Xuejing Liu et al.NeurIPS 2024 · 22 citations
- KnowCoder: Coding Structured Knowledge into LLMs for Universal Information ExtractionZixuan Li, Yutao Zeng, Yuxin Zuo, Weicheng Ren et al.ACL 2024 · 19 citations
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen et al.WWW 2025 · 16 citations
- Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language ModelsQingkai Min, Qipeng Guo, Xiangkun Hu, Songfang Huang et al.ACL 2024 · 11 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Translation and Fusion Improves Cross-lingual Information ExtractionYang Chen, Vedaant Shah, Alan RitterACL 2025
- Guideline Learning for In-Context Information ExtractionChaoxu Pang, Yixuan Cao, Qiang Ding, Ping LuoEMNLP 2023 · 12 citations
- ADELIE: Aligning Large Language Models on Information ExtractionYunjia Qi, Hao Peng, Xiaozhi Wang, Bin Xu et al.EMNLP 2024 · 8 citations
- GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity RecognitionShizhou Huang, Bo Xu, Yang Yu, Changqun Li et al.AAAI 2025 · 1 citation
- CodeIE: Large Code Generation Models are Better Few-Shot Information ExtractorsPeng Li, Tianxiang Sun, Qiong Tang, Hang Yan et al.ACL 2023 · 41 citations
