CalligraphicOCR for Chinese Calligraphy Recognition
Xiaoyi Bao, Zhongqing Wang, Jinghang Gu, Chu-Ren Huang
摘要
With thousand years of history, calligraphy serve as one of the representative symbols of Chinese culture. Increasing works try to digitize calligraphy by recognizing the context of calligraphy for better preservation and propagation. However, previous works stick to isolated single character recognition, not only requires unpractical manual splitting into characters, but also abandon the enriched context information that could be supplementary. To this end, we construct the pioneering end-to-end calligraphy recognition benchmark dataset, this dataset is challenging due to both the visual variations such as different writing styles, and the textual understanding such as the domain shift in semantics. We further propose CalligraphicOCR (COCR) equipped with calligraphic image augmentation and actionbased corrector targeted at the challenging root of this setting. Experiments demonstrate the advantage of our proposed model over cutting-edge baselines, underscoring the necessity of introducing this new setting, thereby facilitating a solid precondition for protecting and propagating the already scarce resources. The code and data are available at https://github.com/HoraceXIaoyiBao/ COCR-EMNLP2025 * Zhongqing Wang and Jinghang Gu are the corresponding authors (Long absent, I miss you deeply. Summer is serene, how fare you? Summoned by duty in old age, I cannot stay. A humble gift of rice conveys my regard. Take care.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Efficient OCR for Building a Diverse Digital HistoryJacob Carlson, Tom Bryan, Melissa DellACL 2024 · 被引用 6 次
- OCR Post Correction for Endangered Language TextsShruti Rijhwani, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 被引用 1 次
- OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by learning to unfoldMohamed Yousef, Tom E. BishopCVPR 2020
- Revisiting Classical Chinese Event Extraction with Ancient Literature InformationXiaoyi Bao, Zhongqing Wang, Jinghang Gu, Chu-Ren HuangACL 2025
相关 Paper
- CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelYuxuan Luo, Jiaqi Tang, Chenyi Huang, Feiyang Hao 等ICCV 2025 · 被引用 4 次
- Reproducing the Past: A Dataset for Benchmarking Inscription RestorationShipeng Zhu, Hui Xue, Na Nie, Chenjie Zhu 等ACM MM 2024 · 被引用 4 次
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 被引用 8 次
- UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese CalligraphyTianshuo Xu, Kai Wang, ZhiFei Chen, Leyi Wu 等ICLR 2026 · 被引用 3 次
- Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-TuningRui Song, Lida Shi, Ruihua Qi, Yingji Li 等ACL 2026 · 被引用 1 次
