Zero-shot Jianzi Recognition as Structured Visual Information Extraction in Open Compositional Symbolic Systems
Zehan Li, Fu Zhang, Zhijun Liu, Jingwei Cheng
Abstract
Guqin (古 琴) Jianzi (減 字) is an open and freely compositional tablature system that encodes performance actions rather than acoustic outcomes. Its automatic recognition remains largely unexplored, as conventional OCR assumes a closed and enumerable glyph set and struggles with Jianzi's unbounded composition and manuscript-level variability. We introduce Zero-shot Jianzi Recognition, which formulates Jianzi recognition as visionto-sequence prediction of canonical component sequences under a zero-shot split. To enable scalable supervision, we construct Synthetic-JZ from aligned online composition metadata. We then synthesize manuscriptlike training images via component-wise style recomposition and manuscript-domain noise modeling, and fine-tune a VLM for end-toend component sequence recognition. At inference time, a lightweight legality-guided correction module re-ranks decoding candidates, suppressing structural hallucinations without modifying the backbone. Experiments on two benchmarks show that our method achieves 63.02% sequence accuracy on Real-JZ, our manually annotated realworld Jianzi benchmark, surpassing Gemini-3-Pro by 35.11%. This result highlights the feasibility of reliable automated Jianzi recognition and its potential for large-scale digitization of historical Guqin Jianzi Pu manuscripts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d71bee0a-2ce3-424f-8491-b7e2650ef828Builds on10
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui et al.AAAI 2023 · 607 citations
- FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningZhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang et al.AAAI 2024 · 90 citations
- Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS AligningHaiyang Yu, Xiaocong Wang, Bin Li, Xiangyang XueICCV 2023 · 43 citations
- We Can Do More to Save Guqin: Design and Evaluate Interactive Systems to Make Guqin More Accessible to the General PublicMinjing Yu, Meng Zhang, Chun Yu, Xiaoguang Ma et al.CHI 2021 · 23 citations
- Deciphering Oracle Bone Language with Diffusion ModelsHaisu Guan, Huanxin Yang, Xinyu Wang, Shengwei Han et al.ACL 2024 · 11 citations
Related papers
- End-to-End Structured Information Extraction from Mixed-Script Documents in Open Compositional Symbol SystemsZehan Li, Fu Zhang, Zhijun Liu, Jingwei ChengKDD 2026
- Toward Zero-shot Character Recognition: A Gold Standard Dataset with Radical-level AnnotationsXiaolei Diao, Daqian Shi, Jian Li, Lida Shi et al.ACM MM 2023 · 11 citations
- BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li et al.ACL 2026
- CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelYuxuan Luo, Jiaqi Tang, Chenyi Huang, Feiyang Hao et al.ICCV 2025 · 4 citations
- Realistic Training Data Generation and Rule Enhanced Decoding in LLM for NameGuessYikuan Xia, Jiazun Chen, Sujian Li, Jun GaoEMNLP 2025
