PAGED: A Benchmark for Procedural Graphs Extraction from Documents
Weihong Du, Wenrui Liao, Hongru Liang, Wenqiang Lei
Abstract
Automatic extraction of procedural graphs from documents creates a low-cost way for users to easily understand a complex procedure by skimming visual graphs. Despite the progress in recent studies, it remains unanswered: whether the existing studies have well solved this task (Q1) and whether the emerging large language models (LLMs) can bring new opportunities to this task (Q2). To this end, we propose a new benchmark PAGED, equipped with a large high-quality dataset and standard evaluations. It investigates five state-of-the-art baselines, revealing that they fail to extract optimal procedural graphs well because of their heavy reliance on hand-written rules and limited available data. We further involve three advanced LLMs in PAGED and enhance them with a novel self-refine strategy. The results point out the advantages of LLMs in identifying textual elements and their gaps in building logical structures. We hope PAGED can serve as a major landmark for automatic procedural graph extraction and the investigations in PAGED can provide valuable insights into the research on logical reasoning among non-sequential elements. The code and dataset are available in https://github.com/SCUNLP/PAGED .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 811bf210-84d8-4f19-a6ed-7007b386245eCited by top-tier papers3
- Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMsChen Yang, Ruping Xu, Ruizhe Li, Bin Cao et al.ACL 2026 · 1 citation
- NL Schedule: Evaluate Multitask Scheduling Capability of Large Language ModelsWenrui Liao, Weihong Du, Yi Li, Hongru Liang et al.ACL 2026
- A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing TaskMashiro Toyooka, Kiyoharu Aizawa, Yoko YamakataACM MM 2025
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Improving Coherence and Consistency in Neural Sequence Models with Dual-System, Neuro-Symbolic ReasoningMaxwell I. Nye, Michael Henry Tessler, Joshua B. Tenenbaum, Brenden M. LakeNeurIPS 2021 · 151 citations
Related papers
- Evaluating LLMs on Large-Scale Graph Property Estimation via Random WalksSunil Kumar Maurya, Xin LiuACL 2026
- GraphSkill: Documentation-Guided Agentic Hierarchical Retrieval-Augmented Coding for Complex Graph ReasoningFali Wang, Chenglin Weng, Xianren Zhang, Siyuan Hong et al.KDD 2026
- LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical StudyDongil Yang, Minjin Kim, Sunghwan Kim, Beong-woo Kwak et al.ACL 2025 · 8 citations
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
- Graph-enhanced Large Language Models in Asynchronous Plan ReasoningFangru Lin, Emanuele La Malfa, Valentin Hofmann, Elle Michelle Yang et al.ICML 2024 · 33 citations
