ActionIE: Action Extraction from Scientific Literature with Programming Languages
Xianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong, Tingfeng Luo, Qirong Ho, Hao Peng, Heng Ji, Jiawei Han
Abstract
Extraction of experimental procedures from human language in scientific literature and patents into actionable sequences in robotics language holds immense significance in scientific domains. Such an action extraction task is particularly challenging given the intricate details and context-dependent nature of the instructions, especially in fields like chemistry where reproducibility is paramount. In this paper, we introduce ACTIONIE, a method that leverages Large Language Models (LLMs) to bridge this divide by converting actions written in natural language into executable Python code. This enables us to capture the entities of interest, and the relationship between each action, given the features of Programming Languages. Utilizing linguistic cues identified by frequent patterns, ActionIE provides an improved mechanism to discern entities of interest. While our method is broadly applicable, we exemplify its power in the domain of chemical literature, wherein we focus on extracting experimental procedures for chemical synthesis. The code generated by our method can be easily transformed into robotics language which is in high demand in scientific fields. Comprehensive experiments demonstrate the superiority of our method. In addition, we propose a graph-based metric to more accurately reflect the precision of extraction. We also develop a dataset to address the scarcity of scientific literature occurred in existing datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40bb9b7c-8c4a-4e93-9ad4-d65d3d940543Cited by top-tier papers2
- ChemReason-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental ChemistryJinwei Zhang, Xucheng Liang, Yu Zhang, Ruijie Yu et al.ACL 2026
- ChemAgent: Self-updating Memories in Large Language Models Improves Chemical ReasoningXiangru Tang, Tianyu Hu, Muyang Ye, Yanjun Shao et al.ICLR 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang et al.ICML 2024 · 443 citations
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang et al.ICML 2024 · 436 citations
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim et al.EMNLP 2022 · 285 citations
- Unlimiformer: Long-Range Transformers with Unlimited Length InputAmanda Bertsch, Uri Alon, Graham Neubig, Matthew GormleyNeurIPS 2023 · 176 citations
Related papers
- ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated DataYu Zhang, Ruijie Yu, Jidong Tian, Feng Zhu et al.ACL 2025 · 3 citations
- PAGED: A Benchmark for Procedural Graphs Extraction from DocumentsWeihong Du, Wenrui Liao, Hongru Liang, Wenqiang LeiACL 2024 · 4 citations
- ReactGPT: Understanding of Chemical Reactions via In-Context TuningZhe Chen, Zhe Fang, Wenhao Tian, Zhaoguang Long et al.AAAI 2025 · 11 citations
- BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in BiologyOdhran O'Donoghue, Aleksandar Shtedritski, John Ginger, Ralph Abboud et al.EMNLP 2023 · 13 citations
- Grounding LLMs in Scientific Discovery via Embodied ActionsBo Zhang, Jinfeng Zhou, Yuxuan Chen, Jianing Yin et al.ICML 2026
