RDLNet: A Novel and Accurate Real-world Document Localization Method
Yaqiang Wu, Zhen Xu, Yong Duan, Yanlai Wu, Qinghua Zheng, Hui Li, Xiaochen Hu, Lianwen Jin
摘要
The increasing use of smartphones for capturing documents in various real-world conditions has underscored the need for robust document localization technologies. Current challenges in this domain include handling diverse document types, complex backgrounds, and varying photographic conditions such as low contrast and occlusion. However, there currently are no publicly available datasets containing these complex scenarios and few methods demonstrate their capabilities on these complex scenes. To address these issues, we create a new comprehensive real-world document localization benchmark dataset which contains the complex scenarios mentioned above and propose a novel Real-world Document Localization Network (RDLNet) for locating targeted documents in the wild. The RDLNet consists of an innovative light-SAM encoder and a masked attention decoder. Utilizing light-SAM encoder, the RDLNet transfers the mighty generalization capability of SAM to the document localization task. In the decoding stage, the RDLNet exploits the masked attention and object query method to efficiently output the triple-branch predictions consisting of corner point coordinates, instance-level segmentation area and categories of different documents without extra post-processing. We compare the performance of RDLNet with other state-of-the-art approaches for real-world document localization on multiple benchmarks, the results of which reveal that the RDLNet remarkably outperforms contemporary methods, demonstrating its superiority in terms of both accuracy and practicability.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Query-driven Generative Network for Document Information Extraction in the WildHaoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang 等ACM MM 2022 · 被引用 13 次
- Learning to Detect Specular Highlights from Real-world ImagesGang Fu, Qing Zhang, Qifeng Lin, Lei Zhu 等ACM MM 2020 · 被引用 45 次
- Fourier Document Restoration for Robust Document Dewarping and RecognitionChuhui Xue, Zichen Tian, Fangneng Zhan, Shijian Lu 等CVPR 2022 · 被引用 37 次
- Learning From Documents in the Wild to Improve Document UnwarpingKe Ma, Sagnik Das, Zhixin Shu, Dimitris SamarasSIGGRAPH 2022 · 被引用 39 次
- Towards Generalized Physical Occlusion Detection On DocumentsYiang Zhu, Haoyue Wang, Zhenxing Qian, Sheng Li 等ACM MM 2025
