Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild
Jiaxin Zhang, Canjie Luo, Lianwen Jin, Fengjun Guo, Kai Ding
摘要
Camera-captured document images usually suffer from perspective and geometric deformations. It is of great value to rectify them when considering poor visual aesthetics and the deteriorated performance of OCR systems. Recent learning-based methods intensively focus on the accurately cropped document image. However, this might not be sufficient for overcoming practical challenges, including document images either with large marginal regions or without margins. Due to this impracticality, users struggle to crop documents precisely when they encounter large marginal regions. Simultaneously, dewarping images without margins is still an insurmountable problem. To the best of our knowledge, there is still no complete and effective pipeline for rectifying document images in the wild. To address this issue, we propose a novel approach called Marior (Margin Removal and Iterative Content Rectification). Marior follows a progressive strategy to iteratively improve the dewarping quality and readability in a coarse-to-fine manner. Specifically, we divide the pipeline into two modules: margin removal module (MRM) and iterative content rectification module (ICRM). First, we predict the segmentation mask of the input image to remove the margin, thereby obtaining a preliminary result. Then we refine the image further by producing dense displacement flows to achieve content-aware rectification. We determine the number of refinement iterations adaptively. Experiments demonstrate the state-of-the-art performance of our method on public benchmarks. The resources are available at https://github.com/ZZZHANG-jx/Marior for further comparison.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual SlimmingJiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie 等AAAI 2025 · 被引用 39 次
- ForCenNet: Foreground-Centric Network for Document Image RectificationPeng Cai, Qiang Li, Kaicheng Yang, Dong Guo 等ICCV 2025 · 被引用 1 次
- DocRes: A Generalist Model Toward Unifying Document Image Restoration TasksJiaxin Zhang, Dezhi Peng, Chongyu Liu, Peirong Zhang 等CVPR 2024
- D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image DewarpingHeng Li, Xiangping Wu, Qingcai ChenCVPR 2026
- M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout AnalysisHiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang 等CVPR 2023
它引用的顶会 Paper3
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- DewarpNet: Single-Image Document Unwarping With Stacked 3D and 2D Regression NetworksSagnik Das, Ke Ma, Zhixin Shu, Dimitris Samaras 等ICCV 2019 · 被引用 97 次
- DocTr: Document Image Transformer for Geometric Unwarping and Illumination CorrectionHao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng 等ACM MM 2021 · 被引用 66 次
相关 Paper
- Revisiting Document Image Dewarping by Grid RegularizationXiangwei Jiang, Rujiao Long, Nan Xue, Zhibo Yang 等CVPR 2022 · 被引用 39 次
- Foreground and Text-lines Aware Document Image RectificationHeng Li, Xiangping Wu, Qingcai Chen, Qianjin XiangICCV 2023 · 被引用 21 次
- Document Registration: Towards Automated Labeling of Pixel-Level Alignment Between Warped-Flat DocumentsWeiguang Zhang, Qiufeng Wang, Kaizhu Huang, Xiaowei Huang 等ACM MM 2024 · 被引用 1 次
- DocDiff: Document Enhancement via Residual Diffusion ModelsZongyuan Yang, Baolin Liu, Yongping Xiong, Lan Yi 等ACM MM 2023 · 被引用 55 次
- End-to-end Piece-wise Unwarping of Document ImagesSagnik Das, Kunwar Yashraj Singh, Jon Wu, Erhan Bas 等ICCV 2021 · 被引用 41 次
