Knowledge-Enhanced Historical Document Segmentation and Recognition
En-Hao Gao, Yu-Xuan Huang, Wen-Chao Hu, Xin-Hao Zhu, Wang-Zhou Dai
摘要
Optical Character Recognition (OCR) of historical document images remains a challenging task because of the distorted input images, extensive number of uncommon characters, and the scarcity of labeled data, which impedes modern deep learning-based OCR techniques from achieving good recognition accuracy. Meanwhile, there exists a substantial amount of expert knowledge that can be utilized in this task. However, such knowledge is usually complicated and could only be accurately expressed with formal languages such as first-order logic (FOL), which is difficult to be directly integrated into deep learning models. This paper proposes KESAR, a novel Knowledge-Enhanced Document Segmentation And Recognition method for historical document images based on the Abductive Learning (ABL) framework. The segmentation and recognition models are enhanced by incorporating background knowledge for character extraction and prediction, followed by an efficient joint optimization of both models. We validate the effectiveness of KESAR on historical document datasets. The experimental results demonstrate that our method can simultaneously utilize knowledge-driven reasoning and data-driven learning, which outperforms the current state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Efficient Rectification of Neuro-Symbolic Reasoning Inconsistencies by Abductive ReflectionWen-Chao Hu, Wang-Zhou Dai, Yuan Jiang, Zhi-Hua ZhouAAAI 2025 · 被引用 14 次
- Ambiguity-Aware Abductive LearningHao-Yuan He, Hui Sun, Zheng Xie, Ming LiICML 2024 · 被引用 6 次
- A learnability analysis on neuro-symbolic learningHao-Yuan He, Ming LiNeurIPS 2025 · 被引用 3 次
- Curriculum Abductive LearningWen-Chao Hu, Qi-Jie Li, Lin-Han Jia, Cunjing Ge 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper9
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang 等ICCV 2019 · 被引用 490 次
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo 等AAAI 2020 · 被引用 289 次
- Neural-Symbolic Integration: A Compositional PerspectiveEfthymia Tsamoura, Timothy M. Hospedales, Loizos MichaelAAAI 2021 · 被引用 85 次
- Fast Abductive Learning by Similarity-based Consistency OptimizationYu-Xuan Huang, Wang-Zhou Dai, Le-Wen Cai, Stephen H. Muggleton 等NeurIPS 2021 · 被引用 42 次
- Injecting Logical Constraints into Neural Networks via Straight-Through EstimatorsZhun Yang, Joohyung Lee, Chiyoun ParkICML 2022 · 被引用 26 次
相关 Paper
- Integrating Deep Learning with Logic Fusion for Information ExtractionWenya Wang, Sinno Jialin PanAAAI 2020 · 被引用 55 次
- LogicSeg: Parsing Visual Semantics with Neural Logic Learning and ReasoningLiulei Li, Wenguan Wang, Yang YiICCV 2023 · 被引用 52 次
- Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan 等ACL 2025 · 被引用 5 次
- Revisiting Document-Level Relation Extraction with Context-Guided Link PredictionMonika Jain, Raghava Mutharaju, Ramakanth Kavuluru, Kuldeep SinghAAAI 2024 · 被引用 17 次
- A Holistic Approach for Answering Logical Queries on Knowledge GraphsYuhan Wu, Yuanyuan Xu, Xuemin Lin, Wenjie ZhangICDE 2023 · 被引用 6 次
