DocTr: Document Image Transformer for Geometric Unwarping and Illumination Correction
Hao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng, Houqiang Li
Abstract
In this work, we propose a new framework, called Document Image Transformer (DocTr), to address the issue of geometry and illumination distortion of the document images. Specifically, DocTr consists of a geometric unwarping transformer and an illumination correction transformer. By setting a set of learned query embedding, the geometric unwarping transformer captures the global context of the document image by self-attention mechanism and decodes the pixel-wise displacement solution to correct the geometric distortion. After geometric unwarping, our illumination correction transformer further removes the shading artifacts to improve the visual quality and OCR accuracy. Extensive evaluations are conducted on several datasets, and superior results are reported against the state-of-the-art methods. Remarkably, our DocTr achieves Character Error Rate (CER), a absolute improvement over the state-of-the-art methods. Moreover, it also shows high efficiency on running time and parameter count.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7480ecb-8707-4efe-8e18-6fd65e8a11daCited by top-tier papers13
- Revisiting Document Image Dewarping by Grid RegularizationXiangwei Jiang, Rujiao Long, Nan Xue, Zhibo Yang et al.CVPR 2022 · 39 citations
- Learning From Documents in the Wild to Improve Document UnwarpingKe Ma, Sagnik Das, Zhixin Shu, Dimitris SamarasSIGGRAPH 2022 · 39 citations
- Fourier Document Restoration for Robust Document Dewarping and RecognitionChuhui Xue, Zichen Tian, Fangneng Zhan, Shijian Lu et al.CVPR 2022 · 37 citations
- Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the WildJiaxin Zhang, Canjie Luo, Lianwen Jin, Fengjun Guo et al.ACM MM 2022 · 25 citations
- Foreground and Text-lines Aware Document Image RectificationHeng Li, Xiangping Wu, Qingcai Chen, Qianjin XiangICCV 2023 · 21 citations
Builds on9
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- DewarpNet: Single-Image Document Unwarping With Stacked 3D and 2D Regression NetworksSagnik Das, Ke Ma, Zhixin Shu, Dimitris Samaras et al.ICCV 2019 · 97 citations
- Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution MatchingJingjun Liang, Ruichen Li, Qin JinACM MM 2020 · 67 citations
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang et al.ACM MM 2020 · 66 citations
- Transformer-based Label Set Generation for Multi-modal Multi-label Emotion DetectionXincheng Ju, Dong Zhang, Junhui Li, Guodong ZhouACM MM 2020 · 64 citations
Related papers
- End-to-end Piece-wise Unwarping of Document ImagesSagnik Das, Kunwar Yashraj Singh, Jon Wu, Erhan Bas et al.ICCV 2021 · 41 citations
- ForCenNet: Foreground-Centric Network for Document Image RectificationPeng Cai, Qiang Li, Kaicheng Yang, Dong Guo et al.ICCV 2025 · 1 citation
- DiT: Self-supervised Pre-training for Document Image TransformerJunlong Li, Yiheng Xu, Tengchao Lv, Lei Cui et al.ACM MM 2022 · 184 citations
- UniDoc: Unified Pretraining Framework for Document UnderstandingJiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao et al.NeurIPS 2021 · 118 citations
- UDoc-GAN: Unpaired Document Illumination Correction with Background Light PriorYonghui Wang, Wengang Zhou, Zhenbo Lu, Houqiang LiACM MM 2022 · 15 citations
