A Scalable Framework for Table of Contents Extraction from Complex ESG Annual Reports
Xinyu Wang, Lin Gui, Yulan He
Abstract
Table of contents (ToC) extraction centres on structuring documents in a hierarchical manner. In this paper, we propose a new dataset, ESGDoc, comprising 1,093 ESG annual reports from 563 companies spanning from 2001 to 2022 . These reports pose significant challenges due to their diverse structures and extensive length. To address these challenges, we propose a new framework for Toc extraction, consisting of three steps: (1) Constructing an initial tree of text blocks based on reading order and font sizes; (2) Modelling each tree node (or text block) independently by considering its contextual information captured in node-centric subtree; (3) Modifying the original tree by taking appropriate action on each tree node (Keep, Delete, or Move). This construction-modellingmodification (CMM) process offers several benefits. It eliminates the need for pairwise modelling of section headings as in previous approaches, making document segmentation practically feasible. By incorporating structured information, each section heading can leverage both local and long-distance context relevant to itself. Experimental results show that our approach outperforms the previous state-of-theart baseline with a fraction of running time. Our framework proves its scalability by effectively handling documents of any length. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f69a8cd-6ca1-4857-aa11-50d527d04f97Cited by top-tier papers1
Ask how each one uses itBuilds on6
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot et al.ACL 2022 · 90 citations
- XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document UnderstandingZhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan et al.CVPR 2022 · 82 citations
- FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat et al.ACL 2023 · 7 citations
Related papers
- Graph-based Document Structure AnalysisYufan Chen, Ruiping Liu, Junwei Zheng, Di Wen et al.ICLR 2025
- Multi-Document Event Extraction Using Large and Small Language ModelsQingkai Min, Zitian Qu, Qipeng Guo, Xiangkun Hu et al.EMNLP 2025 · 1 citation
- ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language ModelsJiani Guo, Zuchao Li, Jie Wu, Qianren Wang et al.EMNLP 2025
- Document-Level Multi-Event Extraction with Event Proxy Nodes and Hausdorff Distance MinimizationXinyu Wang, Lin Gui, Yulan HeACL 2023 · 10 citations
- A Compare-and-contrast Multistage Pipeline for Uncovering Financial Signals in Financial ReportsJia-Huei Ju, Yu-Shiang Huang, Cheng-Wei Lin, Che Lin et al.ACL 2023 · 1 citation
