ActionIE: Action Extraction from Scientific Literature with Programming Languages
Xianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong, Tingfeng Luo, Qirong Ho, Hao Peng, Heng Ji, Jiawei Han
摘要
Extraction of experimental procedures from human language in scientific literature and patents into actionable sequences in robotics language holds immense significance in scientific domains. Such an action extraction task is particularly challenging given the intricate details and context-dependent nature of the instructions, especially in fields like chemistry where reproducibility is paramount. In this paper, we introduce ACTIONIE, a method that leverages Large Language Models (LLMs) to bridge this divide by converting actions written in natural language into executable Python code. This enables us to capture the entities of interest, and the relationship between each action, given the features of Programming Languages. Utilizing linguistic cues identified by frequent patterns, ActionIE provides an improved mechanism to discern entities of interest. While our method is broadly applicable, we exemplify its power in the domain of chemical literature, wherein we focus on extracting experimental procedures for chemical synthesis. The code generated by our method can be easily transformed into robotics language which is in high demand in scientific fields. Comprehensive experiments demonstrate the superiority of our method. In addition, we propose a graph-based metric to more accurately reflect the precision of extraction. We also develop a dataset to address the scarcity of scientific literature occurred in existing datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ChemReason-Bench: Benchmarking Large Language Models for Procedural Reasoning in Experimental ChemistryJinwei Zhang, Xucheng Liang, Yu Zhang, Ruijie Yu 等ACL 2026
- ChemAgent: Self-updating Memories in Large Language Models Improves Chemical ReasoningXiangru Tang, Tianyu Hu, Muyang Ye, Yanjun Shao 等ICLR 2025
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language ModelsAndy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang 等ICML 2024 · 被引用 443 次
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang 等ICML 2024 · 被引用 436 次
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim 等EMNLP 2022 · 被引用 285 次
- Unlimiformer: Long-Range Transformers with Unlimited Length InputAmanda Bertsch, Uri Alon, Graham Neubig, Matthew GormleyNeurIPS 2023 · 被引用 176 次
相关 Paper
- ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated DataYu Zhang, Ruijie Yu, Jidong Tian, Feng Zhu 等ACL 2025 · 被引用 3 次
- PAGED: A Benchmark for Procedural Graphs Extraction from DocumentsWeihong Du, Wenrui Liao, Hongru Liang, Wenqiang LeiACL 2024 · 被引用 4 次
- ReactGPT: Understanding of Chemical Reactions via In-Context TuningZhe Chen, Zhe Fang, Wenhao Tian, Zhaoguang Long 等AAAI 2025 · 被引用 11 次
- BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in BiologyOdhran O'Donoghue, Aleksandar Shtedritski, John Ginger, Ralph Abboud 等EMNLP 2023 · 被引用 13 次
- Grounding LLMs in Scientific Discovery via Embodied ActionsBo Zhang, Jinfeng Zhou, Yuxuan Chen, Jianing Yin 等ICML 2026
