ACL2026
Zero-shot Jianzi Recognition as Structured Visual Information Extraction in Open Compositional Symbolic Systems
Zehan Li, Fu Zhang, Zhijun Liu, Jingwei Cheng
摘要
Guqin (古 琴) Jianzi (減 字) is an open and freely compositional tablature system that encodes performance actions rather than acoustic outcomes. Its automatic recognition remains largely unexplored, as conventional OCR assumes a closed and enumerable glyph set and struggles with Jianzi's unbounded composition and manuscript-level variability. We introduce Zero-shot Jianzi Recognition, which formulates Jianzi recognition as visionto-sequence prediction of canonical component sequences under a zero-shot split. To enable scalable supervision, we construct Synthetic-JZ from aligned online composition metadata. We then synthesize manuscriptlike training images via component-wise style recomposition and manuscript-domain noise modeling, and fine-tune a VLM for end-toend component sequence recognition. At inference time, a lightweight legality-guided correction module re-ranks decoding candidates, suppressing structural hallucinations without modifying the backbone. Experiments on two benchmarks show that our method achieves 63.02% sequence accuracy on Real-JZ, our manually annotated realworld Jianzi benchmark, surpassing Gemini-3-Pro by 35.11%. This result highlights the feasibility of reliable automated Jianzi recognition and its potential for large-scale digitization of historical Guqin Jianzi Pu manuscripts.