Zero-shot Jianzi Recognition as Structured Visual Information Extraction in Open Compositional Symbolic Systems
Zehan Li, Fu Zhang, Zhijun Liu, Jingwei Cheng
摘要
Guqin (古 琴) Jianzi (減 字) is an open and freely compositional tablature system that encodes performance actions rather than acoustic outcomes. Its automatic recognition remains largely unexplored, as conventional OCR assumes a closed and enumerable glyph set and struggles with Jianzi's unbounded composition and manuscript-level variability. We introduce Zero-shot Jianzi Recognition, which formulates Jianzi recognition as visionto-sequence prediction of canonical component sequences under a zero-shot split. To enable scalable supervision, we construct Synthetic-JZ from aligned online composition metadata. We then synthesize manuscriptlike training images via component-wise style recomposition and manuscript-domain noise modeling, and fine-tune a VLM for end-toend component sequence recognition. At inference time, a lightweight legality-guided correction module re-ranks decoding candidates, suppressing structural hallucinations without modifying the backbone. Experiments on two benchmarks show that our method achieves 63.02% sequence accuracy on Real-JZ, our manually annotated realworld Jianzi benchmark, surpassing Gemini-3-Pro by 35.11%. This result highlights the feasibility of reliable automated Jianzi recognition and its potential for large-scale digitization of historical Guqin Jianzi Pu manuscripts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsMinghao Li, Tengchao Lv, Jingye Chen, Lei Cui 等AAAI 2023 · 被引用 607 次
- FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningZhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang 等AAAI 2024 · 被引用 90 次
- Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS AligningHaiyang Yu, Xiaocong Wang, Bin Li, Xiangyang XueICCV 2023 · 被引用 43 次
- We Can Do More to Save Guqin: Design and Evaluate Interactive Systems to Make Guqin More Accessible to the General PublicMinjing Yu, Meng Zhang, Chun Yu, Xiaoguang Ma 等CHI 2021 · 被引用 23 次
- Deciphering Oracle Bone Language with Diffusion ModelsHaisu Guan, Huanxin Yang, Xinyu Wang, Shengwei Han 等ACL 2024 · 被引用 11 次
相关 Paper
- End-to-End Structured Information Extraction from Mixed-Script Documents in Open Compositional Symbol SystemsZehan Li, Fu Zhang, Zhijun Liu, Jingwei ChengKDD 2026
- Toward Zero-shot Character Recognition: A Gold Standard Dataset with Radical-level AnnotationsXiaolei Diao, Daqian Shi, Jian Li, Lida Shi 等ACM MM 2023 · 被引用 11 次
- BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li 等ACL 2026
- CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelYuxuan Luo, Jiaqi Tang, Chenyi Huang, Feiyang Hao 等ICCV 2025 · 被引用 4 次
- Realistic Training Data Generation and Rule Enhanced Decoding in LLM for NameGuessYikuan Xia, Jiazun Chen, Sujian Li, Jun GaoEMNLP 2025
