CalligraphicOCR for Chinese Calligraphy Recognition
Xiaoyi Bao, Zhongqing Wang, Jinghang Gu, Chu-Ren Huang
Abstract
With thousand years of history, calligraphy serve as one of the representative symbols of Chinese culture. Increasing works try to digitize calligraphy by recognizing the context of calligraphy for better preservation and propagation. However, previous works stick to isolated single character recognition, not only requires unpractical manual splitting into characters, but also abandon the enriched context information that could be supplementary. To this end, we construct the pioneering end-to-end calligraphy recognition benchmark dataset, this dataset is challenging due to both the visual variations such as different writing styles, and the textual understanding such as the domain shift in semantics. We further propose CalligraphicOCR (COCR) equipped with calligraphic image augmentation and actionbased corrector targeted at the challenging root of this setting. Experiments demonstrate the advantage of our proposed model over cutting-edge baselines, underscoring the necessity of introducing this new setting, thereby facilitating a solid precondition for protecting and propagating the already scarce resources. The code and data are available at https://github.com/HoraceXIaoyiBao/ COCR-EMNLP2025 * Zhongqing Wang and Jinghang Gu are the corresponding authors (Long absent, I miss you deeply. Summer is serene, how fare you? Summoned by duty in old age, I cannot stay. A humble gift of rice conveys my regard. Take care.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 934390ca-c5dc-4cad-818c-df6854cae09eCited by top-tier papers1
Ask how each one uses itBuilds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Efficient OCR for Building a Diverse Digital HistoryJacob Carlson, Tom Bryan, Melissa DellACL 2024 · 6 citations
- OCR Post Correction for Endangered Language TextsShruti Rijhwani, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 1 citation
- OrigamiNet: Weakly-Supervised, Segmentation-Free, One-Step, Full Page Text Recognition by learning to unfoldMohamed Yousef, Tom E. BishopCVPR 2020
- Revisiting Classical Chinese Event Extraction with Ancient Literature InformationXiaoyi Bao, Zhongqing Wang, Jinghang Gu, Chu-Ren HuangACL 2025
Related papers
- CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelYuxuan Luo, Jiaqi Tang, Chenyi Huang, Feiyang Hao et al.ICCV 2025 · 4 citations
- Reproducing the Past: A Dataset for Benchmarking Inscription RestorationShipeng Zhu, Hui Xue, Na Nie, Chenjie Zhu et al.ACM MM 2024 · 4 citations
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 8 citations
- UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese CalligraphyTianshuo Xu, Kai Wang, ZhiFei Chen, Leyi Wu et al.ICLR 2026 · 3 citations
- Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-TuningRui Song, Lida Shi, Ruihua Qi, Yingji Li et al.ACL 2026 · 1 citation
