FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction
Chen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat, Vincent Perot, Guolong Su, Xiang Zhang, Kihyuk Sohn, Nikolay Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long
摘要
The recent advent of self-supervised pretraining techniques has led to a surge in the use of multimodal learning in form document understanding. However, existing approaches that extend the mask language modeling to other modalities require careful multitask tuning, complex reconstruction target designs, or additional pre-training data. In Form-NetV2, we introduce a centralized multimodal graph contrastive learning strategy to unify self-supervised pre-training for all modalities in one loss. The graph contrastive objective maximizes the agreement of multimodal representations, providing a natural interplay for all modalities without special customization. In addition, we extract image features within the bounding box that joins a pair of tokens connected by a graph edge, capturing more targeted visual cues without loading a sophisticated and separately pre-trained image embedder. FormNetV2 establishes new state-of-theart performance on FUNSD, CORD, SROIE and Payment benchmarks with a more compact model size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Scalable Framework for Table of Contents Extraction from Complex ESG Annual ReportsXinyu Wang, Lin Gui, Yulan HeEMNLP 2023 · 被引用 4 次
- EMC2: Efficient MCMC Negative Sampling for Contrastive Learning with Global ConvergenceChung-Yiu Yau, Hoi-To Wai, Parameswaran Raman, Soumajyoti Sarkar 等ICML 2024 · 被引用 3 次
- Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual SegmentsAniket Bhattacharyya, Anurag Tripathi, Ujjal Das, Archan Karmakar 等ACL 2025
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- Contrastive Multi-View Representation Learning on GraphsKaveh Hassani, Amir Hosein Khas AhmadiICML 2020 · 被引用 1,663 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
相关 Paper
- StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-trainingYuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang 等ICLR 2023 · 被引用 18 次
- LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui 等ACL 2021
- FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot 等ACL 2022 · 被引用 90 次
- UniDoc: Unified Pretraining Framework for Document UnderstandingJiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao 等NeurIPS 2021 · 被引用 118 次
- DocFormerv2: Local Features for Document UnderstandingSrikar Appalaraju, Peng Tang, Qi Dong, Nishant Sankaran 等AAAI 2024 · 被引用 68 次
