CodeIE: Large Code Generation Models are Better Few-Shot Information Extractors
Peng Li, Tianxiang Sun, Qiong Tang, Hang Yan, Yuanbin Wu, Xuanjing Huang, Xipeng Qiu
Abstract
Large language models (LLMs) pre-trained on massive corpora have demonstrated impressive few-shot learning ability on many NLP tasks. A common practice is to recast the task into a text-to-text format such that generative LLMs of natural language (NL-LLMs) like GPT-3 can be prompted to solve it. However, it is nontrivial to perform information extraction (IE) tasks with NL-LLMs since the output of the IE task is usually structured and therefore is hard to be converted into plain text. In this paper, we propose to recast the structured output in the form of code instead of natural language and utilize generative LLMs of code (Code-LLMs) such as Codex to perform IE tasks, in particular, named entity recognition and relation extraction. In contrast to NL-LLMs, we show that Code-LLMs can be well-aligned with these IE tasks by designing code-style prompts and formulating these IE tasks as code generation tasks. Experiment results on seven benchmarks show that our method consistently outperforms fine-tuning moderate-size pre-trained models specially designed for IE tasks (e.g., UIE) and prompting NL-LLMs under few-shot settings. We further conduct a series of in-depth analyses to demonstrate the merits of leveraging Code-LLMs for IE tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6359d6ec-3b80-4209-b5dd-8bbc3aab9111Cited by top-tier papers19
- GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionOscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle et al.ICLR 2024 · 168 citations
- Code-Style In-Context Learning for Knowledge-Based Question AnsweringZhijie Nie, Richong Zhang, Zhongyuan Wang, Xudong LiuAAAI 2024 · 24 citations
- KnowCoder: Coding Structured Knowledge into LLMs for Universal Information ExtractionZixuan Li, Yutao Zeng, Yuxin Zuo, Weicheng Ren et al.ACL 2024 · 19 citations
- QUEST: Query Optimization in Unstructured Document AnalysisZhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin et al.VLDB 2025 · 9 citations
- Improving Natural Language Understanding for LLMs via Large-Scale Instruction SynthesisLin Yuan, Jun Xu, Honghao Gui, Mengshu Sun et al.AAAI 2025 · 3 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim et al.EMNLP 2022 · 285 citations
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang et al.EMNLP 2022 · 103 citations
Related papers
- Generating Diverse Training Samples for Relation Extraction with Large Language ModelsZexuan Li, Hongliang Dai, Piji LiACL 2025
- ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information ExtractionJiabang He, Lei Wang, Yi Hu, Ning Liu et al.ICCV 2023 · 61 citations
- ADELIE: Aligning Large Language Models on Information ExtractionYunjia Qi, Hao Peng, Xiaozhi Wang, Bin Xu et al.EMNLP 2024 · 8 citations
- Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMsHaritz Puerto, Martin Tutek, Somak Aditya, Xiaodan Zhu et al.EMNLP 2024 · 4 citations
- OneNet: A Fine-Tuning Free Framework for Few-Shot Entity Linking via Large Language Model PromptingXukai Liu, Ye Liu, Kai Zhang, Kehang Wang et al.EMNLP 2024 · 7 citations
