End-to-End Structured Information Extraction from Mixed-Script Documents in Open Compositional Symbol Systems
Zehan Li, Fu Zhang, Zhijun Liu, Jingwei Cheng
摘要
Structurally extracting information from mixed-script documents that interleave standard text with open, compositional symbol systems is challenging for both optical character recognition (OCR) and vision–language models (VLMs). This difficulty is epitomized by Jianzi Pu—the ancient Guqin tablature. Unlike closed-set scripts, Jianzi glyphs are formed via infinite compositional rules and are densely integrated with Hanzi text, demanding a model that can simultaneously perform script discrimination, layout parsing, and structural transcription. We propose JZ-Tab, the first framework dedicated to the automated recognition of Jianzi Pu, which functions as an end-to-end structured visual information extraction system for mixed-script documents. Unlike traditional pipelines, JZ-Tab generates layout-aware markup directly from full-page images, bypassing the need for pre-segmentation. Specifically, to overcome the total absence of large-scale annotated datasets, we develop a novel, scalable synthetic-to-real pipeline that constructs layout-consistent pages from canonicalized glyph inventories. Furthermore, to capture the unique action-oriented semantics of the tablature, we introduce music-structured generation, injecting sequential regularities derived from symbolic music logic into the learning process. Finally, we train a VLM for direct page-to-markup generation. Evaluated zero-shot on authentic historical manuscripts, JZ-Tab improves F1 by +40.10 over the strongest generic VLM baselines, highlighting its potential for large-scale automated digitization of historical Guqin manuscripts and open, compositional symbol systems.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Zero-shot Jianzi Recognition as Structured Visual Information Extraction in Open Compositional Symbolic SystemsZehan Li, Fu Zhang, Zhijun Liu, Jingwei ChengACL 2026
- BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li 等ACL 2026
- Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge AugmentationJianing Zhang, Runan Li, Honglin Pang, Ding Xia 等ACL 2026 · 被引用 1 次
- UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese CalligraphyTianshuo Xu, Kai Wang, ZhiFei Chen, Leyi Wu 等ICLR 2026 · 被引用 3 次
- CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelYuxuan Luo, Jiaqi Tang, Chenyi Huang, Feiyang Hao 等ICCV 2025 · 被引用 4 次
