InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling
Xiaolei Diao, Zhihan Zhou, Lida Shi, Ting Wang, Ruihua Qi, Daqian Shi, Hao Xu
Abstract
Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical texts. First, the scarcity of historical language samples renders unsupervised learning approaches based on large text corpora highly inefficient, hindering effective pre-training. Moreover, due to the considerable temporal gap and complex evolution of ancient scripts, the absence of comprehensive character encoding schemes limits the digitization and computational processing of ancient texts, particularly in early Chinese writing. To address these challenges, we introduce InteChar, a unified and extensible character list that integrates unencoded oracle bone characters with traditional and modern Chinese. InteChar enables consistent digitization and representation of historical texts, providing a foundation for robust modeling of ancient scripts. To evaluate the effectiveness of InteChar, we construct the Oracle Corpus Set (OracleCS), an ancient Chinese corpus that combines expert-annotated samples with LLM-assisted data augmentation, centered on Chinese oracle bone inscriptions. Extensive experiments show that models trained with InteChar on Or-acleCS achieve substantial improvements across various historical language understanding tasks, confirming the effectiveness of our approach and establishing a solid foundation for future research in ancient Chinese NLP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f54efd9a-b9e3-4d0c-ae03-4dc7322467b4Builds on5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- CharFormer: A Glyph Fusion based Attentive Framework for High-precision Character Image DenoisingDaqian Shi, Xiaolei Diao, Lida Shi, Hao Tang et al.ACM MM 2022 · 32 citations
- RCRN: Real-world Character Image Restoration Network via Skeleton ExtractionDaqian Shi, Xiaolei Diao, Hao Tang, Xiaomin Li et al.ACM MM 2022 · 25 citations
- Deciphering Oracle Bone Language with Diffusion ModelsHaisu Guan, Huanxin Yang, Xinyu Wang, Shengwei Han et al.ACL 2024 · 11 citations
- Toward Zero-shot Character Recognition: A Gold Standard Dataset with Radical-level AnnotationsXiaolei Diao, Daqian Shi, Jian Li, Lida Shi et al.ACM MM 2023 · 11 citations
Related papers
- OraclePoints: A Hybrid Neural Representation for Oracle CharacterRunhua Jiang, Yongge Liu, Boyuan Zhang, Xu Chen et al.ACM MM 2023 · 9 citations
- V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and MeRunqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu et al.ACL 2025 · 8 citations
- Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge AugmentationJianing Zhang, Runan Li, Honglin Pang, Ding Xia et al.ACL 2026 · 1 citation
- OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?Zijian Chen, Tingzhu Chen, Wenjun Zhang, Guangtao ZhaiICLR 2025 · 3 citations
- AGTGAN: Unpaired Image Translation for Photographic Ancient Character GenerationHongxiang Huang, Daihui Yang, Gang Dai, Zhen Han et al.ACM MM 2022 · 31 citations
