Instruct and Extract: Instruction Tuning for On-Demand Information Extraction
Yizhu Jiao, Ming Zhong, Sha Li, Ruining Zhao, Siru Ouyang, Heng Ji, Jiawei Han
Abstract
Large language models with instructionfollowing capabilities open the door to a wider group of users. However, when it comes to information extraction -a classic task in natural language processing -most task-specific systems cannot align well with long-tail ad hoc extraction use cases for non-expert users. To address this, we propose a novel paradigm, termed On-Demand Information Extraction, to fulfill the personalized demands of real-world users. Our task aims to follow the instructions to extract the desired content from the associated text and present it in a structured tabular format. The table headers can either be userspecified or inferred contextually by the model. To facilitate research in this emerging area, we present a benchmark named INSTRUCTIE, inclusive of both automatically generated training data, as well as the human-annotated test set. Building on INSTRUCTIE, we further develop an On-Demand Information Extractor, ODIE. Comprehensive evaluations on our benchmark reveal that ODIE substantially outperforms the existing open-source models of similar size. Our code and dataset are released on https://github.com/yzjiao/On-Demand-IE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 330bb75e-ce8f-4bca-a989-199484d6bc54Cited by top-tier papers14
- The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT InteractionsSiru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong et al.EMNLP 2023 · 10 citations
- QUEST: Query Optimization in Unstructured Document AnalysisZhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin et al.VLDB 2025 · 9 citations
- ADELIE: Aligning Large Language Models on Information ExtractionYunjia Qi, Hao Peng, Xiaozhi Wang, Bin Xu et al.EMNLP 2024 · 8 citations
- Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMsZhuowen Liang, Xiaotian Lin, Zhengxuan Zhang, Yuyu Luo et al.ICLR 2026 · 6 citations
- Scaling Knowledge Graph Construction through Synthetic Data Generation and DistillationPrafulla Kumar Choubey, Xin Su, Man Luo, XIANGYU PENG et al.ICLR 2026 · 5 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
Related papers
- GoLLIE: Annotation Guidelines improve Zero-Shot Information-ExtractionOscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle et al.ICLR 2024 · 168 citations
- NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured TablesLanrui Wang, Mingyu Zheng, Hongyin Tang, Zheng Lin et al.NeurIPS 2025 · 16 citations
- FollowTable: A Benchmark for Instruction-Following Table RetrievalRihui Jin, Yuchen Lu, Ting Zhang, Jun Wang et al.SIGIR 2026
- Instruction Induction: From Few Examples to Natural Language Task DescriptionsOr Honovich, Uri Shaham, Samuel R. Bowman, Omer LevyACL 2023 · 48 citations
- Ins-DetCLIP: Aligning Detection Model to Follow Human-Language InstructionRenjie Pi, Lewei Yao, Jianhua Han, Xiaodan Liang et al.ICLR 2024 · 5 citations
