A Scalable Framework for Table of Contents Extraction from Complex ESG Annual Reports
Xinyu Wang, Lin Gui, Yulan He
摘要
Table of contents (ToC) extraction centres on structuring documents in a hierarchical manner. In this paper, we propose a new dataset, ESGDoc, comprising 1,093 ESG annual reports from 563 companies spanning from 2001 to 2022 . These reports pose significant challenges due to their diverse structures and extensive length. To address these challenges, we propose a new framework for Toc extraction, consisting of three steps: (1) Constructing an initial tree of text blocks based on reading order and font sizes; (2) Modelling each tree node (or text block) independently by considering its contextual information captured in node-centric subtree; (3) Modifying the original tree by taking appropriate action on each tree node (Keep, Delete, or Move). This construction-modellingmodification (CMM) process offers several benefits. It eliminates the need for pairwise modelling of section headings as in previous approaches, making document segmentation practically feasible. By incorporating structured information, each section heading can leverage both local and long-distance context relevant to itself. Experimental results show that our approach outperforms the previous state-of-theart baseline with a fraction of running time. Our framework proves its scalability by effectively handling documents of any length. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot 等ACL 2022 · 被引用 90 次
- XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document UnderstandingZhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan 等CVPR 2022 · 被引用 82 次
- FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat 等ACL 2023 · 被引用 7 次
相关 Paper
- Graph-based Document Structure AnalysisYufan Chen, Ruiping Liu, Junwei Zheng, Di Wen 等ICLR 2025
- Multi-Document Event Extraction Using Large and Small Language ModelsQingkai Min, Zitian Qu, Qipeng Guo, Xiangkun Hu 等EMNLP 2025 · 被引用 1 次
- ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language ModelsJiani Guo, Zuchao Li, Jie Wu, Qianren Wang 等EMNLP 2025
- Document-Level Multi-Event Extraction with Event Proxy Nodes and Hausdorff Distance MinimizationXinyu Wang, Lin Gui, Yulan HeACL 2023 · 被引用 10 次
- A Compare-and-contrast Multistage Pipeline for Uncovering Financial Signals in Financial ReportsJia-Huei Ju, Yu-Shiang Huang, Cheng-Wei Lin, Che Lin 等ACL 2023 · 被引用 1 次
