UniDoc: Unified Pretraining Framework for Document Understanding
Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, Tong Sun
摘要
Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabeled document datasets have opened up promising directions towards reducing annotation efforts by training models with self-supervised objectives. However, most of the existing document pretraining methods are still language-dominated. We present UDoc, a new unified pretraining framework for document understanding. UDoc is designed to support most document understanding tasks, extending the Transformer to take multimodal embeddings as input. Each input element is composed of words and visual features from a semantic region of the input document image. An important feature of UDoc is that it learns a generic representation by making use of three self-supervised losses, encouraging the representation to model sentences, learn similarities, and align modalities. Extensive empirical analysis demonstrates that the pretraining procedure learns better joint representations and leads to improvements in downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- Document Understanding Dataset and Evaluation (DUDE)Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz 等ICCV 2023 · 被引用 130 次
- Vision Grid Transformer for Document Layout AnalysisCheng Da, Chuwei Luo, Qi Zheng, Cong YaoICCV 2023 · 被引用 63 次
- OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionJianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu 等CVPR 2024 · 被引用 29 次
- Relational Representation Learning in Visually-Rich DocumentsXin Li, Yan Zheng, Yiqing Hu, Haoyu Cao 等ACM MM 2022 · 被引用 7 次
它引用的顶会 Paper8
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
相关 Paper
- SelfDoc: Self-Supervised Document Representation LearningPeizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu 等CVPR 2021
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie 等ICCV 2021 · 被引用 392 次
- Unifying Vision, Text, and Layout for Universal Document ProcessingZineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang 等CVPR 2023
- StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-trainingYuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang 等ICLR 2023 · 被引用 18 次
- DocFormerv2: Local Features for Document UnderstandingSrikar Appalaraju, Peng Tang, Qi Dong, Nishant Sankaran 等AAAI 2024 · 被引用 68 次
