DocTr: Document Transformer for Structured Information Extraction in Documents
Haofu Liao, Aruni RoyChowdhury, Weijian Li, Ankan Bansal, Yuting Zhang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan
Abstract
We present a new formulation for structured information extraction (SIE) from visually rich documents. We address the limitations of existing IOB tagging and graph-based formulations, which are either overly reliant on the correct ordering of input text or struggle with decoding a complex graph. Instead, motivated by anchor-based object detectors in computer vision, we represent an entity as an anchor word and a bounding box, and represent entity linking as the association between anchor words. This is more robust to text ordering, and maintains a compact graph for entity linking. The formulation motivates us to introduce 1) a Document Transformer (DocTr) that aims at detecting and associating entity bounding boxes in visually rich documents, and 2) a simple pre-training strategy that helps learn entity detection in the context of language. Evaluations on three SIE benchmarks show the effectiveness of the proposed formulation, and the overall approach outperforms existing solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bea3c131-4ee5-4f1a-a080-b18357e08510Cited by top-tier papers5
- Hierarchical Visual Feature Aggregation for OCR-Free Document UnderstandingJaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung HanNeurIPS 2024 · 19 citations
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMsAli Faraz, Akash, Shaharukh Khan, Raja Kolla et al.ICLR 2026 · 9 citations
- PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair ExtractionZening Lin, Jiapeng Wang, Teng Li, Wenhui Liao et al.ACM MM 2024 · 6 citations
- Modeling Layout Reading Order as Ordering Relations for Visually-rich Document UnderstandingChong Zhang, Yi Tu, Yixi Zhao, Chenshu Yuan et al.EMNLP 2024 · 4 citations
- Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature RestorationYuxuan Zhou, Baole Wei, Xingjian Hu, Haowei Chen et al.ICML 2026
Builds on12
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
- BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from DocumentsTeakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang et al.AAAI 2022 · 186 citations
Related papers
- StrucTexT: Structured Text Understanding with Multi-Modal TransformersYulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin et al.ACM MM 2021 · 124 citations
- UniDoc: Unified Pretraining Framework for Document UnderstandingJiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao et al.NeurIPS 2021 · 118 citations
- UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich DocumentsYi Tu, Chong Zhang, Ya Guo, Huan Chen et al.ACM MM 2024 · 2 citations
- Improving Information Extraction from Visually Rich Documents using Visual Span RepresentationsRitesh Sarkhel, Arnab NandiVLDB 2021 · 17 citations
- DocTr: Document Image Transformer for Geometric Unwarping and Illumination CorrectionHao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng et al.ACM MM 2021 · 66 citations
